VLDB 2026 Research / reviewers in the wild / expert
Costantino Grana
dblp:02/5570
· DBLP profile ↗
80ranked-venue papers
19as first author
19since 2021 · last 2026
0000-0002-4792-2358ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 15 first-author · 13 since 2021Artificial intelligence and machine learning · 35 · 6 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling 8B Bitwise Autoregressive Image Generation on Edge GPUs
Enrico Vezzali, Federico Bolelli, Costantino Grana, Luca Benini, Yawei Li 0001 |
ICPR (12) | 3 |
| 2026 | Multi-structure segmentation in CBCT volumes: The ToothFairy2 challengeabstractCone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT. Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana |
Medical Image Anal. | 36 |
| 2026 | A State-of-the-Art Review With Code About Connected Components Labeling on GPUsabstractThis article is about Connected Components Labeling (CCL) algorithms developed for GPU accelerators. The task itself is employed in many modern image-processing pipelines and represents a fundamental step in different scenarios, whenever object recognition is required. For this reason, a strong effort in the development of many different proposals devoted to improving algorithm performance using different kinds of hardware accelerators has been made. This paper focuses on GPU-based algorithmic solutions published in the last two decades, highlighting their distinctive traits and the improvements they leverage. The state-of-the-art review proposed is equipped with the source code, which allows to straightforwardly reproduce all the algorithms in different experimental settings. A comprehensive evaluation on multiple environments is also provided, including different operating systems, compilers, and GPUs. Our assessments are performed by means of several tests, including real-case images and synthetically generated ones, highlighting the strengths and weaknesses of each proposal. Overall, the experimental results revealed that block-based oriented algorithms outperform all the other algorithmic solutions on both 2D images and 3D volumes, regardless of the selected environment. Federico Bolelli, Stefano Allegretti, Luca Lumetti, Costantino Grana |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | A Deep-Learning-Based Method for Real-Time Barcode Segmentation on Edge CPUs
Enrico Vezzali, Lorenzo Vorabbi, Costantino Grana, Federico Bolelli |
CAIP (1) | 3 |
| 2025 | Segmenting Maxillofacial Structures in CBCT VolumesabstractCone-Beam computed tomography (CBCT) is a standard imaging modality in orofacial and dental practices, providing essential 3D volumetric imaging of anatomical structures, including jawbones, teeth, sinuses, and neurovascular canals. Accurately segmenting these structures is fundamental to numerous clinical applications, such as surgical planning and implant placement. However, manual segmentation of CBCT scans is time-intensive and requires expert input, creating a demand for automated solutions through deep learning. Effective development of such algorithms relies on access to large, well-annotated datasets, yet current datasets are often privately stored or limited in scope and considered structures, especially concerning 3D annotations. This paper proposes ToothFairy2, a comprehensive, publicly accessible CBCT dataset with voxel-level 3D annotations of 42 distinct classes corresponding to maxillofacial structures. We validate the dataset by benchmarking state-of-the-art neural network models, including convolutional, transformer-based, and hybrid Mamba-Based architectures, to evaluate segmentation performance across complex anatomical regions. Our work also explores adaptations to the nnU-Net framework to optimize multi-class segmentation for maxillofacial anatomy. The proposed dataset provides a fundamental resource for advancing maxillofacial segmentation and supports future research in automated 3D image analysis in digital dentistry. Federico Bolelli, Kevin Marchesini, Niels van Nistelrooij, Luca Lumetti, Vittorio Pipoli, Elisa Ficarra, Shankeeth Vinayahalingam, Costantino Grana |
CVPR | 8 |
| 2025 | MISSRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
Vittorio Pipoli, Alessia Saporita, Federico Bolelli, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
ICCV | 6 |
| 2025 | Mosaic-SR: An Adaptive Multi-Step Super-Resolution Method For Low-Resolution 2d BarcodesabstractQR and Datamatrix codes are widely used in warehouse logistics and high-speed production pipelines. Still, distant or small barcodes often yield low-pixel-density images that are hard to read. Conventional solutions rely on costly hardware or enhanced lighting, raising expenses and potentially reducing depth of field. We propose Mosaic-SR, a multi-step, adaptive super-resolution (SR) method that devotes more computation to barcode regions than uniform backgrounds. For each patch, it predicts an uncertainty value to decide how many refinement steps are required. Our experiments show that Mosaic-SR surpasses state-of-the-art SR models on 2D barcode images, achieving higher PSNR and decoding rates in less time. All code and trained models are publicly available at https://github.com/Henvezz95/mosaic-sr Enrico Vezzali, Lorenzo Vorabbi, Costantino Grana, Federico Bolelli |
ICIP | 3 |
| 2025 | U-Net Transplant: The Role of Pre-training for Model Merging in 3D Medical Segmentation
Luca Lumetti, Giacomo Capitani, Elisa Ficarra, Simone Calderara, Costantino Grana, Angelo Porrello, Federico Bolelli |
MICCAI (16) | 5 |
| 2025 | IM-Fuse: A Mamba-Based Fusion Block for Brain Tumor Segmentation with Incomplete Modalities
Vittorio Pipoli, Alessia Saporita, Kevin Marchesini, Costantino Grana, Elisa Ficarra, Federico Bolelli |
MICCAI (8) | 4 |
| 2025 | Semantically Conditioned Prompts for Visual Recognition Under Missing Modality ScenariosabstractThis paper tackles the domain of multimodal prompting for visual recognition, specifically when dealing with missing modalities through multimodal Transformers. It presents two main contributions: (i) we introduce a novel prompt learning module which is designed to produce sample-specific prompts and (ii) we show that modalityagnostic prompts can effectively adjust to diverse missing modality scenarios. Our model, termed SCP, exploits the semantic representation of available modalities to query a learnable memory bank, which allows the generation of prompts based on the semantics of the input. Notably, SCP distinguishes itself from existing methodologies for its capacity of self-adjusting to both the missing modality scenario and the semantic context of the input, without prior knowledge about the specific missing modality and the number of modalities. Through extensive experiments, we show the effectiveness of the proposed prompt learning framework and demonstrate enhanced performance and robustness across a spectrum of missing modality cases. Our source code is available at https://github.com/vittoriopipoli/SCP_WACV2025. Vittorio Pipoli, Federico Bolelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara, Elisa Ficarra |
WACV | 6 |
| 2025 | State-of-the-art review and benchmarking of barcode localization methodsabstractBarcodes, despite their long history, remain an essential technology in supply chain management. In addition, barcodes have found extensive use in industrial engineering, particularly in warehouse automation, component tracking, and robot guidance. To detect a barcode in an image, multiple algorithms have been proposed in the literature, with a significant increase of interest in the topic since the rise of deep learning. However, research in the field suffers from many limitations, including the scarcity of public datasets and code implementations which hinders the reproducibility and reliability of published results. For this reason, we developed “BarBeR” (Barcode Benchmark Repository), a benchmark designed for testing and comparing barcode detection algorithms. This benchmark includes the code implementation of various detection algorithms for barcodes, along with a suite of useful metrics. Among the supported localization methods, there are multiple deep-learning detection models, that will be used to assess the recent contributions of Artificial Intelligence to this field. In addition, we provide a large, annotated dataset of 8 748 barcode images, combining multiple public barcode datasets with standardized annotation formats for both detection and segmentation tasks. Finally, we provide a thorough summary of the history and literature on barcode localization and share the results obtained from running the benchmark on our dataset, offering valuable insights into the performance of different algorithms when applied to real-world problems. Enrico Vezzali, Federico Bolelli, Stefano Santi, Costantino Grana |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Segmenting the Inferior Alveolar Canal in CBCTs Volumes: The ToothFairy ChallengeabstractIn recent years, several algorithms have been developed for the segmentation of the Inferior Alveolar Canal (IAC) in Cone-Beam Computed Tomography (CBCT) scans. However, the availability of public datasets in this domain is limited, resulting in a lack of comparative evaluation studies on a common benchmark. To address this scientific gap and encourage deep learning research in the field, the ToothFairy challenge was organized within the MICCAI 2023 conference. In this context, a public dataset was released to also serve as a benchmark for future research. The dataset comprises 443 CBCT scans, with voxel-level annotations of the IAC available for 153 of them, making it the largest publicly available dataset of its kind. The participants of the challenge were tasked with developing an algorithm to accurately identify the IAC using the 2D and 3D-annotated scans. This paper presents the details of the challenge and the contributions made by the most promising methods proposed by the participants. It represents the first comprehensive comparative evaluation of IAC segmentation methods on a common benchmark dataset, providing insights into the current state-of-the-art algorithms and outlining future research directions. Furthermore, to ensure reproducibility and promote future developments, an open-source repository that collects the implementations of the best submissions was released. Federico Bolelli, Luca Lumetti, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Arrigo Pellacani, Kevin Marchesini, Niels van Nistelrooij, Pieter van Lierop, Tong Xi 0001, Yusheng Liu 0001, Rui Xin 0003, Tao Yang 0037, Lisheng Wang, Haoshen Wang, Chenfan Xu, Zhiming Cui 0001, Marek Wodzinski, Henning Müller, Yannick Kirchhoff, Maximilian Rokuss, Klaus H. Maier-Hein, Jae-Hwan Han, Wan Kim, Hong-Gi Ahn, Tomasz Szczepanski, Michal K. Grzeszczyk, Przemyslaw Korzeniowski, Vicent Caselles, Xavier Paolo Burgos-Artizzu, Ferran Prados, Stefaan Bergé, Bram van Ginneken, Alexandre Anesi, Costantino Grana |
IEEE Trans. Medical Imaging | 34 |
| 2024 | Investigating the ABCDE Rule in Convolutional Neural Networks
Federico Bolelli, Luca Lumetti, Kevin Marchesini, Ettore Candeloro, Costantino Grana |
ICPR (13) | 5 |
| 2024 | Location Matters: Harnessing Spatial Information to Enhance the Segmentation of the Inferior Alveolar Canal in CBCTs
Luca Lumetti, Vittorio Pipoli, Federico Bolelli, Elisa Ficarra, Costantino Grana |
ICPR (28) | 5 |
| 2024 | Identifying Impurities in Liquids of Pharmaceutical Vials
Gabriele Rosati, Kevin Marchesini, Luca Lumetti, Federica Sartori, Beatrice Balboni, Filippo Begarani, Luca Vescovi, Federico Bolelli, Costantino Grana |
ICPR (17) | 9 |
| 2024 | BarBeR: A Barcode Benchmarking Repository
Enrico Vezzali, Federico Bolelli, Stefano Santi, Costantino Grana |
ICPR (17) | 4 |
| 2022 | Improving Segmentation of the Inferior Alveolar Nerve through Deep Label PropagationabstractMany recent works in dentistry and maxillofacial imagery focused on the Inferior Alveolar Nerve (IAN) canal detection. Unfortunately, the small extent of available 3D maxillofacial datasets has strongly limited the performance of deep learning-based techniques. On the other hand, a huge amount of sparsely annotated data is produced every day from the regular procedures in the maxillofacial practice. Despite the amount of sparsely labeled images being significant, the adoption of those data still raises an open problem. Indeed, the deep learning approach frames the presence of dense annotations as a crucial factor. Recent efforts in literature have hence focused on developing label propagation techniques to expand sparse annotations into dense labels. However, the proposed methods proved only marginally effective for the purpose of segmenting the alveolar nerve in CBCT scans. This paper exploits and publicly releases a new 3D densely annotated dataset, through which we are able to train a deep label propagation model which obtains better results than those available in literature. By combining a segmentation model trained on the 3D annotated data and label propagation, we significantly improve the state of the art in the Inferior Alveolar Nerve segmentation. Marco Cipriano, Stefano Allegretti, Federico Bolelli, Federico Pollastri, Costantino Grana |
CVPR | 5 |
| 2022 | One DAG to Rule Them AllabstractIn this paper, we present novel strategies for optimizing the performance of many binary image processing algorithms. These strategies are collected in an open-source framework, GRAPHGEN, that is able to automatically generate optimized C++ source code implementing the desired optimizations. Simply starting from a set of rules, the algorithms introduced with the GRAPHGEN framework can generate decision trees with minimum average path-length, possibly considering image pattern frequencies, apply state prediction and code compression by the use of Directed Rooted Acyclic Graphs (DRAGs). Moreover, the proposed algorithmic solutions allow to combine different optimization techniques and significantly improve performance. Our proposal is showcased on three classical and widely employed algorithms (namely Connected Components Labeling, Thinning, and Contour Tracing). When compared to existing approaches -in 2D and 3D-, implementations using the generated optimal DRAGs perform significantly better than previous state-of-the-art algorithms, both on CPU and GPU. Federico Bolelli, Stefano Allegretti, Costantino Grana |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | A deep analysis on high-resolution dermoscopic image classificationabstractAbstract Convolutional neural networks (CNNs) have been broadly employed in dermoscopic image analysis, mainly as a result of the large amount of data gathered by the International Skin Imaging Collaboration (ISIC). As in many other medical imaging domains, state‐of‐the‐art methods take advantage of architectures developed for other tasks, frequently assuming full transferability between enormous sets of natural images (e.g. ImageNet) and dermoscopic images, which is not always the case. A comprehensive analysis on the effectiveness of state‐of‐the‐art deep learning techniques when applied to dermoscopic image analysis is provided. To achieve this goal, the authors consider several CNNs architectures and analyse how their performance is affected by the size of the network, image resolution, data augmentation process, amount of available data, and model calibration. Moreover, taking advantage of the analysis performed, a novel ensemble method to further increase the classification accuracy is designed. The proposed solution achieved the third best result in the 2019 official ISIC challenge, with an accuracy of 0.593. Federico Pollastri, Mario Parreño, Juan Maroñas Molano, Federico Bolelli, Roberto Paredes, Daniel Ramos-Castro, Costantino Grana |
IET Comput. Vis. | 7 |
| 2020 | Supporting Skin Lesion Diagnosis with Content-Based Image RetrievalabstractIn recent years, many attempts have been dedicated to the creation of automated devices that could assist both expert and beginner dermatologists towards fast and early diagnosis of skin lesions. Tasks such as skin lesion classification and segmentation have been extensively addressed with deep learning algorithms, which in some cases reach a diagnostic accuracy comparable to that of expert physicians. However, the general lack of interpretability and reliability severely hinders the ability of those approaches to actually support dermatologists in the diagnosis process. In this paper a novel skin image retrieval system is presented, which exploits features extracted by Convolutional Neural Networks to gather similar images from a publicly available dataset, in order to assist the diagnosis process of both expert and novice practitioners. In the proposed framework, ResNet-50 is initially trained for the classification of dermoscopic images; then, the feature extraction part is isolated, and an embedding network is built on top of it. The embedding learns an alternative representation, which allows to check image similarity by means of a distance measure. Experimental results reveal that the proposed method is able to select meaningful images, which can effectively boost the classification accuracy of human dermatologists. Stefano Allegretti, Federico Bolelli, Federico Pollastri, Sabrina Longhitano, Giovanni Pellacani, Costantino Grana |
ICPR | 6 |
| 2020 | The DeepHealth Toolkit: A Unified Framework to Boost Biomedical ApplicationsabstractGiven the overwhelming impact of machine learning on the last decade, several libraries and frameworks have been developed in recent years to simplify the design and training of neural networks, providing array-based programming, automatic differentiation and user-friendly access to hardware accelerators. None of those tools, however, was designed with native and transparent support for Cloud Computing or heterogeneous High-Performance Computing (HPC). The DeepHealth Toolkit is an open source Deep Learning toolkit aimed at boosting productivity of data scientists operating in the medical field by providing a unified framework for the distributed training of neural networks, which is able to leverage hybrid HPC and cloud environments in a transparent way for the user. The toolkit is composed of a Computer Vision library, a Deep Learning library, and a front-end for non-expert users; all of the components are focused on the medical domain, but they are general purpose and can be applied to any other field. In this paper, the principles driving the design of the DeepHealth libraries are described, along with details about the implementation and the interaction between the different elements composing the toolkit. Finally, experiments on common benchmarks prove the efficiency of each separate component and of the DeepHealth Toolkit overall. Michele Cancilla, Laura Canalini, Federico Bolelli, Stefano Allegretti, Salvador Carrión-Ponz, Roberto Paredes, Jon Ander Gómez, Simone Leo, Marco Enrico Piras, Luca Pireddu, Asaf Badouh, Santiago Marco-Sola, Lluc Alvarez, Miquel Moretó, Costantino Grana |
ICPR | 15 |
| 2020 | Confidence Calibration for Deep Renal Biopsy Immunofluorescence Image ClassificationabstractWith this work we tackle immunofluorescence classification in renal biopsy, employing state-of-the-art Convolutional Neural Networks. In this setting, the aim of the probabilistic model is to assist an expert practitioner towards identifying the location pattern of antibody deposits within a glomerulus. Since modern neural networks often provide overconfident outputs, we stress the importance of having a reliable prediction, demonstrating that Temperature Scaling (TS), a recently introduced re-calibration technique, can be successfully applied to immunofluorescence classification in renal biopsy. Experimental results demonstrate that the designed model yields good accuracy on the specific task, and that TS is able to provide reliable probabilities, which are highly valuable for such a task given the low inter-rater agreement. Federico Pollastri, Juan Maroñas Molano, Federico Bolelli, Giulia Ligabue, Roberto Paredes, Riccardo Magistroni, Costantino Grana |
ICPR | 7 |
| 2020 | A Heuristic-Based Decision Tree for Connected Components Labeling of 3D VolumesabstractConnected Components Labeling represents a fundamental step for many Computer Vision and Image Processing pipelines. Since the first appearance of the task in the sixties, many algorithmic solutions to optimize the computational load needed to label an image have been proposed. Among them, block-based scan approaches and decision trees revealed to be some of the most valuable strategies. However, due to the cost of the manual construction of optimal decision trees and the computational limitations of automatic strategies employed in the past, the application of blocks and decision trees has been restricted to small masks, and thus to 2D algorithms. With this paper we present a novel heuristic algorithm based on decision tree learning methodology, called Entropy Partitioning Decision Tree (EPDT). It allows to compute near-optimal decision trees for large scan masks. Experimental results demonstrate that algorithms based on the generated decision trees outperform state-of-the-art competitors. Maximilian Söchting, Stefano Allegretti, Federico Bolelli, Costantino Grana |
ICPR | 4 |
| 2020 | Augmenting data with GANs to segment melanoma skin lesions
Federico Pollastri, Federico Bolelli, Roberto Paredes, Costantino Grana |
Multim. Tools Appl. | 4 |
| 2020 | Spaghetti Labeling: Directed Acyclic Graphs for Block-Based Connected Components LabelingabstractConnected Components Labeling is an essential step of many Image Processing and Computer Vision tasks. Since the first proposal of a labeling algorithm, which dates back to the sixties, many approaches have optimized the computational load needed to label an image. In particular, the use of decision forests and state prediction have recently appeared as valuable strategies to improve performance. However, due to the overhead of the manual construction of prediction states and the size of the resulting machine code, the application of these strategies has been restricted to small masks, thus ignoring the benefit of using a block-based approach. In this paper, we combine a block-based mask with state prediction and code compression: the resulting algorithm is modeled as a Directed Rooted Acyclic Graph with multiple entry points, which is automatically generated without manual intervention. When tested on synthetic and real datasets, in comparison with optimized implementations of state-of-the-art algorithms, the proposed approach shows superior performance, surpassing the results obtained by all compared approaches in all settings. Federico Bolelli, Stefano Allegretti, Lorenzo Baraldi 0001, Costantino Grana |
IEEE Trans. Image Process. | 4 |
| 2020 | Optimized Block-Based Algorithms to Label Connected Components on GPUsabstractConnected Components Labeling (CCL) is a crucial step of several image processing and computer vision pipelines. Many efficient sequential strategies exist, among which one of the most effective is the use of a block-based mask to drastically cut the number of memory accesses. In the last decade, aided by the fast development of Graphics Processing Units (GPUs), a lot of data parallel CCL algorithms have been proposed along with sequential ones. Applications that entirely run in GPU can benefit from parallel implementations of CCL that allow to avoid expensive memory transfers between host and device. In this paper, two new eight-connectivity CCL algorithms are proposed, namely Block-based Union Find (BUF) and Block-based Komura Equivalence (BKE). These algorithms optimize existing GPU solutions introducing a block-based approach. Extensions for three-dimensional datasets are also discussed. In order to produce a fair comparison with previously proposed alternatives, YACCLAB, a public CCL benchmarking framework, has been extended and made suitable for evaluating also GPU algorithms. Moreover, three-dimensional datasets have been added to its collection. Experimental results on real cases and synthetically generated datasets demonstrate the superiority of the new proposals with respect to state-of-the-art, both on 2D and 3D scenarios. Stefano Allegretti, Federico Bolelli, Costantino Grana |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | How Does Connected Components Labeling with Decision Trees Perform on GPUs?
Stefano Allegretti, Federico Bolelli, Michele Cancilla, Federico Pollastri, Laura Canalini, Costantino Grana |
CAIP (1) | 6 |
| 2019 | Skin Lesion Segmentation Ensemble with Diverse Training Strategies
Laura Canalini, Federico Pollastri, Federico Bolelli, Michele Cancilla, Stefano Allegretti, Costantino Grana |
CAIP (1) | 6 |
| 2018 | Improving Skin Lesion Segmentation with Generative Adversarial NetworksabstractThis paper proposes a novel strategy that employs Generative Adversarial Networks (GANs) to augment data in the image segmentation field, and a Convolutional-Deconvolutional Neural Network (CDNN) to automatically generate lesion segmentation mask from dermoscopic images. Training the CDNN with our GAN generated data effectively improves the state-of-the-art. Federico Bolelli, Federico Pollastri, Roberto Paredes, Costantino Grana |
CBMS | 4 |
| 2018 | Aligning Text and Document Illustrations: Towards Visually Explainable Digital HumanitiesabstractWhile several approaches to bring vision and language together are emerging, none of them has yet addressed the digital humanities domain, which, nevertheless, is a rich source of visual and textual data. To foster research in this direction, we investigate the learning of visual-semantic embeddings for historical document illustrations, devising both supervised and semi-supervised approaches. We exploit the joint visual-semantic embeddings to automatically align illustrations and textual elements, thus providing an automatic annotation of the visual content of a manuscript. Experiments are performed on the Borso d'Este Holy Bible, one of the most sophisticated illuminated manuscript from the Renaissance, which we manually annotate aligning every illustration with textual commentaries written by experts. Experimental results quantify the domain shift between ordinary visual-semantic datasets and the proposed one, validate the proposed strategies, and devise future works on the same line. Lorenzo Baraldi 0001, Marcella Cornia, Costantino Grana, Rita Cucchiara |
ICPR | 3 |
| 2018 | Connected Components Labeling on DRAGsabstractIn this paper we introduce a new Connected Components Labeling (CCL) algorithm which exploits a novel approach to model decision problems as Directed Acyclic Graphs with a root, which will be called Directed Rooted Acyclic Graphs (DRAGs). This structure supports the use of sets of equivalent actions, as required by CCL, and optimally leverages these equivalences to reduce the number of nodes (decision points). The advantage of this representation is that a DRAG, differently from decision trees usually exploited by the state-of-the-art algorithms, will contain only the minimum number of nodes required to reach the leaf corresponding to a set of condition values. This combines the benefits of using binary decision trees with a reduction of the machine code size. Experiments show a consistent improvement of the execution time when the model is applied to CCL. Federico Bolelli, Lorenzo Baraldi 0001, Michele Cancilla, Costantino Grana |
ICPR | 4 |
| 2018 | Optimizing GPU-Based Connected Components Labeling AlgorithmsabstractConnected Components Labeling (CCL) is a fundamental image processing technique, widely used in various application areas. Computational throughput of Graphical Processing Units (GPUs) makes them eligible for such a kind of algorithms. In the last decade, many approaches to compute CCL on GPUs have been proposed. Unfortunately, most of them have focused on 4-way connectivity neglecting the importance of 8-way connectivity. This paper aims to extend state-of-the-art GPU-based algorithms from 4 to 8-way connectivity and to improve them with additional optimizations. Experimental results revealed the effectiveness of the proposed strategies. Stefano Allegretti, Federico Bolelli, Michele Cancilla, Costantino Grana |
IPAS | 4 |
| 2018 | A Hierarchical Quasi-Recurrent approach to Video CaptioningabstractVideo captioning has picked up a considerable attention thanks to the ability of Recurrent Neural Networks to extrapolate an encoded representation of the input video, and then use it to generate a description. We propose a recurrent encoding approach able to find and exploit the layered design of the video. Differently from the established encoder-decoder procedure, in which a video is repeatedly encoded by a recurrent layer, we employ revised Quasi-Recurrent Neural Networks. We further extend their basic cell with a boundary detector in order to recognize discontinuous segments boundaries and likewise correct the temporal connections of the encoding layer accordingly. Experiments, on the Montreal Video Annotation dataset, demonstrate that our approach can find suitable levels of representation of the input information, while reducing the computational requirements. Federico Bolelli, Lorenzo Baraldi 0001, Federico Pollastri, Costantino Grana |
IPAS | 4 |
| 2017 | Hierarchical Boundary-Aware Neural Encoder for Video CaptioningabstractThe use of Recurrent Neural Networks for video captioning has recently gained a lot of attention, since they can be used both to encode the input video and to generate the corresponding description. In this paper, we present a recurrent video encoding scheme which can discover and leverage the hierarchical structure of the video. Unlike the classical encoder-decoder approach, in which a video is encoded continuously by a recurrent layer, we propose a novel LSTM cell which can identify discontinuity points between frames or segments and modify the temporal connections of the encoding layer accordingly. We evaluate our approach on three large-scale datasets: the Montreal Video Annotation dataset, the MPII Movie Description dataset and the Microsoft Video Description Corpus. Experiments show that our approach can discover appropriate hierarchical representations of input videos and improve the state of the art results on movie description datasets. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
CVPR | 2 |
| 2017 | Segmentation models diversity for object proposals
Marco Manfredi, Costantino Grana, Rita Cucchiara, Arnold W. M. Smeulders |
Comput. Vis. Image Underst. | 2 |
| 2017 | Recognizing and Presenting the Storytelling Video Structure With Deep Multimodal NetworksabstractIn this paper, we propose a novel scene detection algorithm which employs semantic, visual, textual, and audio cues. We also show how the hierarchical decomposition of the storytelling video structure can improve retrieval results presentation with semantically and aesthetically effective thumbnails. Our method is built upon two advancements of the state of the art: first is semantic feature extraction which builds video-specific concept detectors; and second is multimodal feature embedding learning that maps the feature vector of a shot to a space in which the Euclidean distance has task specific semantic properties. The proposed method is able to decompose the video in annotated temporal segments which allow us for a query specific thumbnail extraction. Extensive experiments are performed on different data sets to demonstrate the effectiveness of our algorithm. An in-depth discussion on how to deal with the subjectivity of the task is conducted and a strategy to overcome the problem is suggested. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
IEEE Trans. Multim. | 2 |
| 2017 | Affective level design for a role-playing videogame evaluated by a brain-computer interface and machine learning methods
Fabrizio Balducci, Costantino Grana, Rita Cucchiara |
Vis. Comput. | 2 |
| 2016 | Optimized Connected Components Labeling with Pixel Prediction
Costantino Grana, Lorenzo Baraldi 0001, Federico Bolelli |
ACIVS | 1 |
| 2016 | Historical document digitization through layout analysis and deep content classificationabstractDocument layout segmentation and recognition is an important task in the creation of digitized documents collections, especially when dealing with historical documents. This paper presents an hybrid approach to layout segmentation as well as a strategy to classify document regions, which is applied to the process of digitization of an historical encyclopedia. Our layout analysis method merges a classic top-down approach and a bottom-up classification process based on local geometrical features, while regions are classified by means of features extracted from a Convolutional Neural Network merged in a Random Forest classifier. Experiments are conducted on the first volume of the “Enciclopedia Treccani”, a large dataset containing 999 manually annotated pages from the historical Italian encyclopedia. Andrea Corbelli, Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICPR | 3 |
| 2016 | YACCLAB - Yet Another Connected Components Labeling BenchmarkabstractThe problem of labeling the connected components (CCL) of a binary image is well-defined and several proposals have been presented in the past. Since an exact solution to the problem exists and should be mandatory provided as output, algorithms mainly differ on their execution speed. In this paper, we propose and describe YACCLAB, Yet Another Connected Components Labeling Benchmark. Together with a rich and varied dataset, YACCLAB contains an open source platform to test new proposals and to compare them with publicly available competitors. Textual and graphical outputs are automatically generated for three kinds of test, which analyze the methods from different perspectives. The fairness of the comparisons is guaranteed by running on the same system and over the same datasets. Examples of usage and the corresponding comparisons among state-of-the-art techniques are reported to confirm the potentiality of the benchmark. Costantino Grana, Federico Bolelli, Lorenzo Baraldi 0001, Roberto Vezzani |
ICPR | 1 |
| 2016 | Scene-driven Retrieval in Edited Videos using Aesthetic and Semantic Deep FeaturesabstractThis paper presents a novel retrieval pipeline for video collections, which aims to retrieve the most significant parts of an edited video for a given query, and represent them with thumbnails which are at the same time semantically meaningful and aesthetically remarkable. Videos are first segmented into coherent and story-telling scenes, then a retrieval algorithm based on deep learning is proposed to retrieve the most significant scenes for a textual query. A ranking strategy based on deep features is finally used to tackle the problem of visualizing the best thumbnail. Qualitative and quantitative experiments are conducted on a collection of edited videos to demonstrate the effectiveness of our approach. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICMR | 2 |
| 2016 | A Browsing and Retrieval System for Broadcast Videos using Scene Detection and Automatic AnnotationabstractThis paper presents a novel video access and retrieval system for edited videos. The key element of the proposal is that videos are automatically decomposed into semantically coherent parts (called scenes) to provide a more manageable unit for browsing, tagging and searching. The system features an automatic annotation pipeline, with which videos are tagged by exploiting both the transcript and the video itself. Scenes can also be retrieved with textual queries; the best thumbnail for a query is selected according to both semantics and aesthetics criteria. Lorenzo Baraldi 0001, Costantino Grana, Alberto Messina, Rita Cucchiara |
ACM Multimedia | 2 |
| 2016 | Guest Editorial: Multimedia for Cultural Heritage
Costantino Grana, Giuseppe Serra 0001 |
Multim. Tools Appl. | 1 |
| 2016 | Layout analysis and content enrichment of digitized books
Costantino Grana, Giuseppe Serra 0001, Marco Manfredi, Dalia Coppi, Rita Cucchiara |
Multim. Tools Appl. | 1 |
| 2015 | Shot and Scene Detection via Hierarchical Clustering for Re-using Broadcast Video
Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
CAIP (1) | 2 |
| 2015 | Scene segmentation using temporal clustering for accessing and re-using broadcast videoabstractScene detection is a fundamental tool for allowing effective video browsing and re-using. In this paper we present a model that automatically divides videos into coherent scenes, which is based on a novel combination of local image descriptors and temporal clustering techniques. Experiments are performed to demonstrate the effectiveness of our approach, by comparing our algorithm against two recent proposals for automatic scene segmentation. We also propose improved performance measures that aim to reduce the gap between numerical evaluation and expected results. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ICME | 2 |
| 2015 | A Deep Siamese Network for Scene Detection in Broadcast VideosabstractWe present a model that automatically divides broadcast videos into coherent scenes by learning a distance measure between shots. Experiments are performed to demonstrate the effectiveness of our approach by comparing our algorithm against recent proposals for automatic scene segmentation. We also propose an improved performance measure that aims to reduce the gap between numerical evaluation and expected results, and propose and release a new benchmark dataset. Lorenzo Baraldi 0001, Costantino Grana, Rita Cucchiara |
ACM Multimedia | 2 |
| 2015 | GOLD: Gaussians of Local Descriptors for image representation
Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2014 | Learning superpixel relations for supervised image segmentationabstractIn this paper we propose to extend the well known graph cut segmentation framework by learning superpixel relations and use them to weight superpixel-to-superpixel edges in a superpixel graph. Adjacent superpixel-pairs are analyzed to build an object boundary model, able to discriminate between superpixel-pairs belonging to the same object or placed on the edge between the foreground object and the background. Several superpixel-pair features are investigated and exploited to build a non-linear SVM to learn object boundary appearance. The adoption of this modified graph cut enhances the performance of a previously proposed segmentation method on two publicly available datasets, reaching state-of-the-art results. Marco Manfredi, Costantino Grana, Rita Cucchiara |
ICIP | 2 |
| 2014 | Truncated isotropic principal component classifier for image classificationabstractThis paper reports a novel approach to deal with the problem of Object and Scene recognition extending the traditional Bag of Words approach in two ways. Firstly, a dataset independent method of summarizing local features, based on multivariate Gaussian descriptors, is employed. Secondly, a recently proposed classification technique, particularly suited for high dimensional feature spaces without any dimensionality reduction step, allows to effectively exploit these features. Experiments are performed on two publicly available datasets and demonstrate the effectiveness of our approach when compared to state-of-the-art methods. Alessandro Rozza, Giuseppe Serra 0001, Costantino Grana |
ICIP | 3 |
| 2014 | Learning Graph Cut Energy Functions for Image SegmentationabstractIn this paper we address the task of learning how to segment a particular class of objects, by means of a training set of images and their segmentations. In particular we propose a method to overcome the extremely high training time of a previously proposed solution to this problem, Kernelized Structural Support Vector Machines. We employ a one-class SVM working with joint kernels to robustly learn significant support vectors (representative image-mask pairs) and accordingly weight them to build a suitable energy function for the graph cut framework. We report results obtained on two public datasets and a comparison of training times on different training set sizes. Marco Manfredi, Costantino Grana, Rita Cucchiara |
ICPR | 2 |
| 2014 | Covariance of Covariance Features for Image ClassificationabstractIn this paper we propose a novel image descriptor built by computing the covariance of pixel level features on densely sampled patches and encoding them using their covariance. Appropriate projections to the Euclidean space and feature normalizations are employed in order to provide a strong descriptor usable with linear classifiers. In order to remove border effects, we further enhance the Spatial Pyramid representation with bilinear interpolation. Experimental results conducted on two common datasets for object and texture classification show that the performance of our method is comparable with state of the art techniques, but removing any dataset specific dependency in the feature encoding step. Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara |
ICMR | 2 |
| 2014 | Miniature illustrations retrieval and innovative interaction for digital illuminated manuscripts
Daniele Borghesani, Costantino Grana, Rita Cucchiara |
Multim. Syst. | 2 |
| 2014 | A complete system for garment segmentation and color classification
Marco Manfredi, Costantino Grana, Simone Calderara, Rita Cucchiara |
Mach. Vis. Appl. | 2 |
| 2013 | Modeling local descriptors with multivariate gaussians for object and scene recognitionabstractCommon techniques represent images by quantizing local descriptors and summarizing their distribution in a histogram. In this paper we propose to employ a parametric description and compare its capabilities to histogram based approaches. We use the multivariate Gaussian distribution, applied over the SIFT descriptors, extracted with dense sampling on a spatial pyramid. Every distribution is converted to a high-dimensional descriptor, by concatenating the mean vector and the projection of the covariance matrix on the Euclidean space tangent to the Riemannian manifold. Experiments on Caltech-101 and ImageCLEF2011 are performed using the Stochastic Gradient Descent solver, which allows to deal with large scale datasets and high dimensional feature spaces. Giuseppe Serra 0001, Costantino Grana, Marco Manfredi, Rita Cucchiara |
ACM Multimedia | 2 |
| 2013 | Beyond Bag of Words for Concept Detection and Search of Cultural Heritage Archives
Costantino Grana, Giuseppe Serra 0001, Marco Manfredi, Rita Cucchiara |
SISAP | 1 |
| 2012 | Class-Based Color Bag of Words for Fashion RetrievalabstractColor signatures, histograms and bag of colors are basic and effective strategies for describing the color content of images, for retrieving images by their color appearance or providing color annotation. In some domains, colors assume a specific meaning for users and the color-based classification and retrieval should mirror the initial suggestions given by users in the training set. For instance in fashion world, the names given to the dominant color of a garment or a dress reflect the fashion dictact and not an uniform division of the color space. In this paper we propose a general approach to implement color signature as a trained bag of words, defined on the basis of user defined color classes. The novel Class-based Color Bag of Words is a easy computable bag of words of color, constructed following an approach similar to the Median Cut algorithm, but biased by color distribution in the trained classes. Moreover, to dramatically reduce the computational effort we propose 3D integral histograms, a 3D extension of integral images, easily extensible for many histogram-based signature in 3D color space. Several comparisons in large fashion datasets confirm the discriminant power of this signature. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
ICME | 1 |
| 2012 | 2D images map warping for improved user interaction
Daniele Borghesani, Costantino Grana, Rita Cucchiara |
ICPR | 2 |
| 2012 | Learning non-target items for interesting clothes segmentation in fashion images
Costantino Grana, Simone Calderara, Daniele Borghesani, Rita Cucchiara |
ICPR | 1 |
| 2012 | Veiling Luminance estimation on FPGA-based embedded smart cameraabstractThis paper describes the design and development of a Veiling Luminance estimation system based on the use of a CMOS image sensor, fully implemented on FPGA. The system is composed of the CMOS Image sensor, FPGA, DDR SDRAM, USB controller and SPI (Serial Peripheral Interface) Flash. The FPGA is used to build a system-on-chip integrating a soft processor (Xilinx MicroBlaze) and all the hardware blocks needed to handle the external peripherals and memory. The soft processor is used to handle image acquisition and all computational tasks need to compute the Veiling Luminance value. The advantages of this single chip FPGA implementation include the reduction of the hardware requirements, power consumption, and system complexity. The problem of the high dynamic range images have been addressed with multiple acquisitions at different exposure times. Vignetting, radial distortion and angular weighting, as required by veiling luminance definition, are handled by a single integer look-up table (LUT) access. Results are compared with a state of the art certified instrument. Costantino Grana, Daniele Borghesani, Paolo Santinelli, Rita Cucchiara |
Intelligent Vehicles Symposium | 1 |
| 2012 | Optimal decision trees for local image processing algorithms
Costantino Grana, Manuela Montangero, Daniele Borghesani |
Pattern Recognit. Lett. | 1 |
| 2011 | Feature Space Warping Relevance Feedback with Transductive Learning
Daniele Borghesani, Dalia Coppi, Costantino Grana, Simone Calderara, Rita Cucchiara |
ACIVS | 3 |
| 2011 | Relevance feedback strategies for artistic image collections taggingabstractThis paper provides an analysis on relevance feedback techniques in a multimedia system designed for the interactive exploration and annotation of artistic collections, in particular illuminated manuscripts. The relevance feedback is presented not only as a very effective technique to improve the performance of the system, but also as a clever way to increase the user experience, mixing the interactive surfing through the artistic content with the possibility to gather valuable information from the user, and consequently improving his retrieval satisfaction. We compare a modification of the Mean-Shift Feature Space Warping algorithm, as representative of the standard RF procedures, and a learning-based technique based on transduction, considered in order to overcome some limitation of the previous technique. Experiments are reported regarding the adopted visual features based on covariance matrices. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
ICMR | 1 |
| 2011 | Automatic segmentation of digitalized historical manuscripts
Costantino Grana, Daniele Borghesani, Rita Cucchiara |
Multim. Tools Appl. | 1 |
| 2011 | Probabilistic people tracking with appearance models and occlusion classification: The AD-HOC system
Roberto Vezzani, Costantino Grana, Rita Cucchiara |
Pattern Recognit. Lett. | 2 |
| 2010 | Decision Trees for Fast Thinning AlgorithmsabstractWe propose a new efficient approach for neighborhood exploration, optimized with decision tables and decision trees, suitable for local algorithms in image processing. In this work, it is employed to speed up two widely used thinning techniques. The performance gain is shown over a large freely available dataset of scanned document images. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
ICPR | 1 |
| 2010 | Surfing on artistic documents with visually assisted taggingabstractThis paper describes a complete architecture for the interactive exploration and annotation of artistic collections. In particular the focus is on Renaissance illuminated manuscripts, which typically contain thousands of pictures, used to comment or embellish the manuscript Gothic text. The final aim is to create a human centered multimedia application allowing the non practitioners to enjoy these masterpieces and expert users to share their knowledge. The system is composed by a modern user interface for browsing, surfing and querying, an automatic segmentation module, to ease the initial picture extraction task, and a similarity based retrieval engine, used to provide visually assisted tagging capabilities. A relevance feedback procedure is included to further refine the results. Experiments are reported regarding the adopted visual features based on covariance matrices and the Mean Shift Feature Space Warping relevance feedback. Finally some hints on the user interface for museum installations are discussed. Daniele Borghesani, Costantino Grana, Rita Cucchiara |
ACM Multimedia | 2 |
| 2010 | Rerum novarum: interactive exploration of illuminated manuscriptsabstractThis paper describes an interactive application for the exploration and annotation of illuminated manuscripts, which typically contain thousands of pictures, used to comment or embellish the manuscript Gothic text. The system is composed by a modern user interface for browsing, surfing and querying, an automatic segmentation module, to ease the initial picture extraction task, and a similarity based retrieval engine, used to provide visually assisted tagging capabilities. A relevance feedback procedure is included to further refine the results. Daniele Borghesani, Costantino Grana, Rita Cucchiara |
ACM Multimedia | 2 |
| 2010 | Optimized Block-Based Connected Components Labeling With Decision TreesabstractIn this paper, we define a new paradigm for eight-connection labeling, which employs a general approach to improve neighborhood exploration and minimizes the number of memory accesses. First, we exploit and extend the decision table formalism introducing OR-decision tables, in which multiple alternative actions are managed. An automatic procedure to synthesize the optimal decision tree from the decision table is used, providing the most effective conditions evaluation order. Second, we propose a new scanning technique that moves on a 2 x 2 pixel grid over the image, which is optimized by the automatically generated decision tree. An extensive comparison with the state of art approaches is proposed, both on synthetic and real datasets. The synthetic dataset is composed of different sizes and densities random images, while the real datasets are an artistic image analysis dataset, a document analysis dataset for text detection and recognition, and finally a standard resolution dataset for picture segmentation tasks. The algorithm provides an impressive speedup over the state of the art algorithms. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
IEEE Trans. Image Process. | 1 |
| 2009 | Fast block based connected components labelingabstractIn this paper we present a new optimization technique for the neighborhood computation in connected component labeling focused on images stored in raster scan order. This new technique is based on a 2 × 2 square block analysis of the image, and it exploits the fact that, when using 8-connection, the pixels of a 2 × 2 square are all connected to each other. This implies that they will share the same label at the end of the computation. To prove the effectiveness of our proposal, we show a comprehensive comparison of the most used and advanced connected components labeling techniques presented so far. The tests are conducted on high resolution images obtained from digitized historical manuscripts and a set of transformations is applied in order to show the algorithms behavior at different image resolutions and with a varying number of labels. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
ICIP | 1 |
| 2008 | Describing texture directions with Von Mises distributionsabstractIn this work we describe a new approach for texture characterization. Starting from the autocorrelation matrix an elegant description through a mixture of Von Mises distributions is proposed. A compact 6 valued descriptor is produced for each block and served as input to an SVM classifier. Tests are carried out on high resolution illuminated manuscripts images. Costantino Grana, Daniele Borghesani, Rita Cucchiara |
ICPR | 1 |
| 2007 | Linear Transition Detection as a Unified Shot Detection ApproachabstractIn this paper, we propose an automatic system for video shot segmentation, called linear transition detector, unique for both cuts and linear transitions detection. Comparison with publicly available shot detection systems is reported on different sports (Formula 1 racing, basketball, soccer, and cycling) and TRECVID 2005 results are also reported Costantino Grana, Rita Cucchiara |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | A Semi-Automatic Video Annotation tool with MPEG-7 Content CollectionsabstractIn this work, we present a general purpose system for hierarchical structural segmentation and automatic annotation of video clips, by means of standardized low level features. We propose to automatically extract some prototypes for each class with a context based intra-class clustering. Clips are annotated following the MPEG-7 standard directives to provide easier portability. Results of automatic annotation and semiautomatic metadata creation are provided Roberto Vezzani, Costantino Grana, Daniele Bulgarelli, Rita Cucchiara |
ISM | 2 |
| 2006 | MOM: multimedia ontology manager. A framework for automatic annotation and semantic retrieval of video sequencesabstractEffective usage of multimedia digital libraries has to deal with the problem of building efficient content annotation and retrieval tools. MOM (Multimedia Ontology Manager) is a complete system that allows the creation of multimedia ontologies, supports automatic annotation and creation of extended text (and audio) commentaries of video sequences, and permits complex queries by reasoning on the ontology. Marco Bertini 0001, Alberto Del Bimbo, Carlo Torniai, Rita Cucchiara, Costantino Grana |
ACM Multimedia | 5 |
| 2006 | PEANO: pictorial enriched annotation of videoabstractIn this DEMO, we present a tool set for video digital library management that allows i) structural annotation of edited videos in MPEG-7 by automatically extracting shots and clips; ii) automatic semantic annotation based on perceptual similarity against a taxonomy enriched with pictorial concepts iii) video clip access and hierarchical summarization with stand-alone and web interface iv) access to clips from mobile platform in GPRS-UMTS video-streaming. The tools can be applied in different domain-specific Video Digital Libraries. The main novelty is the possibility to enrich the annotation with pictorial concepts that are added to a textual taxonomy in order to make the automatic annotation process more fast and often effective. The resulting multimedia ontology is described in the MPEG-7 framework. The PEANO (Perceptual Annotation of Video) tool has been tested over video art , sport (Soccer, Olimpic Games 2006, Formula 1) and news clips. Costantino Grana, Roberto Vezzani, Daniele Bulgarelli, Giovanni Gualdi, Rita Cucchiara, Marco Bertini 0001, Carlo Torniai, Alberto Del Bimbo |
ACM Multimedia | 1 |
| 2005 | MPEG-7 Compliant Shot Detection in Sport VideosabstractIn this paper we propose a system for automatic detection of shots in sport videos. Our work covers two main aspects: the first is robust shot detection in presence of fast object motion and camera operations. To this aim we propose a new algorithm, unique for both cuts and linear transitions detection, which only needs the tuning of two parameters. An extended comparison with four transition detection algorithms, representing the state of the art in literature, is reported. Examples with formula 1, basket, soccer and cycling videos are analyzed. The second aspect is an in depth discussion on the annotation of shots and transitions with the MPEG-7 standard. Costantino Grana, Giovanni Tardini, Rita Cucchiara |
ISM | 1 |
| 2005 | Probabilistic posture classification for Human-behavior analysisabstractComputer vision and ubiquitous multimedia access nowadays make feasible the development of a mostly automated system for human-behavior analysis. In this context, our proposal is to analyze human behaviors by classifying the posture of the monitored person and, consequently, detecting corresponding events and alarm situations, like a fall. To this aim, our approach can be divided in two phases: for each frame, the projection histograms (Haritaoglu et al., 1998) of each person are computed and compared with the probabilistic projection maps stored for each posture during the training phase; then, the obtained posture is further validated exploiting the information extracted by a tracking module in order to take into account the reliability of the classification of the first phase. Moreover, the tracking algorithm is used to handle occlusions, making the system particularly robust even in indoors environments. Extensive experimental results demonstrate a promising average accuracy of more than 95% in correctly classifying human postures, even in the case of challenging conditions. Rita Cucchiara, Costantino Grana, Andrea Prati 0001, Roberto Vezzani |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2003 | Detecting Moving Objects, Ghosts, and Shadows in Video StreamsabstractBackground subtraction methods are widely exploited for moving object detection in videos in many applications, such as traffic monitoring, human motion capture, and video surveillance. How to correctly and efficiently model and update the background model and how to deal with shadows are two of the most distinguishing and challenging aspects of such approaches. The article proposes a general-purpose method that combines statistical assumptions with the object-level knowledge of moving objects, apparent objects (ghosts), and shadows acquired in the processing of the previous frames. Pixels belonging to moving objects, ghosts, and shadows are processed differently in order to supply an object-based selective update. The proposed approach exploits color information for both background subtraction and shadow detection to improve object segmentation and background update. The approach proves fast, flexible, and precise in terms of both pixel accuracy and reactivity to background changes. Rita Cucchiara, Costantino Grana, Massimo Piccardi, Andrea Prati 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | A New Algorithm for Border Description of Polarized Light Surface Microscopic Images of Pigmented Skin LesionsabstractThe aim of this study was to provide mathematical descriptors for the border of pigmented skin lesion images and to assess their efficacy for distinction among different lesion groups. New descriptors such as lesion slope and lesion slope regularity are introduced and mathematically defined. A new algorithm based on the Catmull-Rom spline method and the computation of the gray-level gradient of points extracted by interpolation of normal direction on spline points was employed. The efficacy of these new descriptors was tested on a data set of 510 pigmented skin lesions, composed by 85 melanomas and 425 nevi, by employing statistical methods for discrimination between the two populations. Costantino Grana, Giovanni Pellacani, Rita Cucchiara, Stefania Seidenari |
IEEE Trans. Medical Imaging | 1 |
| 2002 | Semantic transcoding for live video serverabstractIn this paper we present transcoding techniques for a video server architecture that enables the user to access live video streams by using different devices with different capabilities. For live videos, annotation methods cannot be exploited. Instead we propose methods of on-the-fly transcoding that adapt the video content with respect to the user resources and the video semantic. Thus we propose an object-based transcoding with classes of relevance (for instance People, Face and Background). To compare the different strategies we propose a metric based on the Weighted Mean Square Error that allows the analysis of different application scenarios by means of a class-wise distortion measure. The obtained results show that the use of semantic can improve the bandwidth to distortion ratio significantly. Rita Cucchiara, Costantino Grana, Andrea Prati 0001 |
ACM Multimedia | 2 |