VLDB 2026 Research / reviewers in the wild / expert
Riccardo Leonardi
dblp:24/4295
· DBLP profile ↗
116ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-0755-1924ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging PerspectiveabstractLearning-based image coding is showing improved compression efficiency, while also offering a novel advantage in enabling computer vision tasks directly within the compressed domain. The latent representation created by deep learning methods inherently contains all visual features, without a computationally expensive synthesis process at the decoder. This paper is an invited extension of a previous solution for JPEG AI compressed domain face detection that adapts a RetinaFace-based detector to operate directly on the latent tensor. In addition to a former single-scale bridging solution, this work provides a novel multi-scale bridging architecture to enable a more effective multi-scale compressed domain face detection. The results show a significant performance gain, improving accuracy up to 20% for detection of tiny faces on the WIDER FACE dataset compared to single-scale bridging, and further narrowing the gap when compared to detection on uncompressed or JPEG AI decoded images. Furthermore, since the computationally expensive decoding step is bypassed and since the bridges consist of lower-complexity networks, the overall processing cost is significantly reduced. Single and multi-scale bridging, respectively, have about 10% and 32% the complexity of applying pixel domain face detection on decoded images. The proposed architecture is expected to be extended to other multiscale sensitive vision tasks, as JPEG AI is not specifically designed for any single downstream application. Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Learning Optimal Linear Block Transform by Rate Distortion MinimizationabstractThe rise of deep learning has spurred advancements in image compression, with end-to-end learned systems gaining traction. However, their adoption in standard frameworks is limited, as they require a major overhaul of existing hardware designed for traditional methods. Moreover, their computational complexity, especially on the decoder side, remains significantly higher than conventional codecs. Consequently, optimizing traditional codecs remains a key research focus. Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
DCC | 4 |
| 2025 | Bridging Compressed Image Latents and Multimodal Large Language ModelsabstractThis paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to modalities (e.g. images) beyond text, but their billion scale hinders deployment on resource-constrained end devices. While cloud-hosted MLLMs could be available, transmitting raw, uncompressed images captured by end devices to the cloud requires an efficient image compression system. To address this, we focus on emerging neural image compression and propose a novel framework with a lightweight transform-neck and a surrogate loss to adapt compressed image latents for MLLM-based vision tasks.
Given the huge scale of MLLMs, our framework excludes the entire downstream MLLM except part of its visual encoder from training our system. This stands out from most existing coding for machine approaches that involve downstream networks in training and thus could be impractical when the networks are MLLMs. The proposed framework is general in that it is applicable to various MLLMs, neural image codecs, and multiple application scenarios, where the neural image codec can be (1) pre-trained for human perception without updating, (2) fully updated for joint human and machine perception, or (3) fully updated for only machine perception.
Extensive experiments on different neural image codecs and various MLLMs show that our method achieves great rate-accuracy performance with much less complexity. Chia-Hao Kao, Cheng Chien, Yu-Jen Tseng, Yi-Hsin Chen, Alessandro Gnutti, Shao-Yuan Lo, Wen-Hsiao Peng, Riccardo Leonardi |
ICLR | 8 |
| 2025 | Rate Distortion Learned Transform For Image Compression
Alessandro Gnutti, Chia-Hao Kao, Wen-Hsiao Peng, Riccardo Leonardi |
PCS | 4 |
| 2025 | Learned Transcoding for Neural Image Compression
Chia-Hao Kao, Andrea Migliorati, Alessandro Gnutti, Enrico Magli, Riccardo Leonardi |
PCS | 5 |
| 2025 | Variable-Size Symmetry-Based Graph Fourier Transforms for Image Compression
Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Pgnn-Based Approach for Robust 3D Light Direction Estimation in Outdoor ImagesabstractEstimating the 3D light direction from 2D outdoor images is a crucial task in computer vision, especially useful in the context of image forgery detection. In this paper, we propose a novel approach that leverages a physics-guided neural network (PGNN) to achieve accurate global 3D light direction estimation. The proposed architecture incorporates an illumination model that enables the network to indirectly learn geometric information and improve the accuracy of the estimated light direction. To evaluate the performance of our proposed approach, we train and test our models on two datasets that were curated for the global light direction estimation. The proposed PGNN method demonstrates superior performance compared to existing state-of-the-art approaches, where direct comparisons were feasible. To further investigate the benefits provided by embedding the physical model in our approach, we conducted an extensive ablation study, demonstrating that the use of the illumination model significantly enhances the accuracy of the light direction estimation compared to a purely data-driven approach. The proposed PGNN is available as open-source software, providing an accessible and useful tool for researchers in computer vision and graphics. Marcello Zanardelli, Mahyar Gohari, Riccardo Leonardi, Sergio Benini, Nicola Adami |
CBMI | 3 |
| 2024 | Lidar Depth Map Guided Image Compression ModelabstractThe incorporation of LiDAR technology into some high-end smartphones has unlocked numerous possibilities across various applications, including photography, image restoration, augmented reality, and more. In this paper, we introduce a novel direction that harnesses LiDAR depth maps to enhance the compression of the corresponding RGB camera images. To the best of our knowledge, this represents the initial exploration in this particular research direction. Specifically, we propose a Transformer-based learned image compression system capable of achieving variable-rate compression using a single model while utilizing the LiDAR depth map as supplementary information for both the encoding and decoding processes. Experimental results demonstrate that integrating LiDAR yields an average PSNR gain of 0.83 dB and an average bitrate reduction of 16% as compared to its absence. Alessandro Gnutti, Stefano Della Fiore, Mattia Savardi, Yi-Hsin Chen, Riccardo Leonardi, Wen-Hsiao Peng |
ICIP | 5 |
| 2024 | JPEG AI Compressed Domain Face DetectionabstractLearning-based image coding has achieved competitive performance in terms of compression efficiency, while also gaining a key advantage in the ability to carry out computer vision tasks directly in the compressed domain. In fact, the latent representation which is generated using deep learning techniques may natively encapsulate all visual features needed for processing tasks, thereby eliminating the need to perform the expensive synthesis transform process at the decoder side. In this paper, it is proposed to perform face detection using the latent code present in the JPEG AI architecture. First, some experiments show how decoded images can be efficiently processed for face detection without retraining, albeit with some performance degradation. Then, for the first time a compressed domain RetinaFace-based detector applied to JPEG AI latent representations is competitively proposed. The performance achieved is comparable to the performance of the original RetinaFace applied to the reconstructed JPEG AI images, while reducing computational complexity since it bypasses the image decoding process. It is expected that this approach might be extended to other vision tasks since the JPEG AI representation format is not tailored specifically for any computer vision task. Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001 |
MMSP | 4 |
| 2024 | Transformer-Based Learned Image Compression for Joint Decoding and DenoisingabstractThis work introduces a Transformer-based image compression system. It has the flexibility to switch between the standard image reconstruction and the denoising reconstruction from a single compressed bitstream. Instead of training separate decoders for these tasks, we incorporate two add-on modules to adapt a pre-trained image decoder from performing the standard image reconstruction to joint decoding and denoising. Our scheme adopts a two-pronged approach. It features a latent refinement module to refine the latent representation of a noisy input image for reconstructing a noise-free image. Additionally, it incorporates an instance-specific prompt generator that adapts the decoding process to improve on the latent refinement. Experimental results show that our method achieves a similar level of denoising quality to training a separate decoder for joint decoding and denoising at the expense of only a modest increase in the decoder's model size and computational complexity. Yi-Hsin Chen, Kuan-Wei Ho, Shiau-Rung Tsai, Guan-Hsun Lin, Alessandro Gnutti, Wen-Hsiao Peng, Riccardo Leonardi |
PCS | 7 |
| 2023 | Image forgery detection: a survey of recent deep-learning approachesabstractAbstract In the last years, due to the availability and easy of use of image editing tools, a large amount of fake and altered images have been produced and spread through the media and the Web. A lot of different approaches have been proposed in order to assess the authenticity of an image and in some cases to localize the altered (forged) areas. In this paper, we conduct a survey of some of the most recent image forgery detection methods that are specifically designed upon Deep Learning (DL) techniques, focusing on commonly found copy-move and splicing attacks. DeepFake generated content is also addressed insofar as its application is aimed at images, achieving the same effect as splicing. This survey is especially timely because deep learning powered techniques appear to be the most relevant right now, since they give the best overall performances on the available benchmark datasets. We discuss the key-aspects of these methods, while also describing the datasets on which they are trained and validated. We also discuss and compare (where possible) their performance. Building upon this analysis, we conclude by addressing possible future research trends and directions, in both deep learning architectural and evaluation approaches, and dataset building for easy methods comparison. Marcello Zanardelli, Fabrizio Guerrini, Riccardo Leonardi, Nicola Adami |
Multim. Tools Appl. | 3 |
| 2022 | Machine learning techniques for MRI feature-based detection of frontotemporal lobar degenerationabstractMaking a diagnosis of neurodegenerative diseases at an early stage is one of the most significant challenges of modern neuroscience. Although this family of diseases remains without a cure, the effectiveness of their medical treatment largely relies on the timing of their detection. For certain groups of diseases, such as Fronto-Temporal Dementia (FTD), trained professionals can effectively reach a correct diagnosis through the visual analysis of Magnetic Resonance Imaging, in its functional (fMRI) or raw (MRI) version. However, this operation is time-consuming and may be subject to personal interpretation. In this paper, we explore the performance of a group of machine learning algorithms to formulate a correct FTD diagnosis, in order to provide medical professionals with a supporting tool. The dataset consists of MRI data acquired on 30 subjects, and the experiments are carried out by investigating different fMRI techniques based on a Multi-Voxel Pattern Analysis (MVPA) approach. The results obtained show high accuracy in identifying FTD in elderly patients when Support Vector Machine and Random Forest techniques are used, with outcomes varying based on the fMRI methods. Tatiana Pilipenko, Alessandro Gnutti, Andrea Silvestri, Ivan Serina, Riccardo Leonardi |
KES | 5 |
| 2021 | Symmetry-Based Graph Fourier Transforms: Are They Optimal For Image Compression?abstractTraditional block-based transforms are based on applying a single transform to all blocks. As an alternative, better performance in image and video processing and representation can be achieved by choosing one among a discrete set of transforms for each block. As an example, our recently proposed set of multiple transforms called Symmetry-Based Graph Fourier Transforms (SBGFTs) have shown good performance in terms of energy compaction, improving HEVC intra coding performance when used to replace the Discrete Cosine Transform (DCT). This paper further explores the performance of the SBGFTs in a multiple transforms, non-linear approximation perspective, by comparing them with two alternative sets of orthogonal transforms, namely, the Karhunen-Loève Transform (KLT) and the Sparse Orthonormal Transform (SOT). Experimental results confirm that SBGFTs achieve superior representation ability in this context as well, suggesting that they could assume a central role in image compression. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 3 |
| 2021 | BS-Net: Learning COVID-19 pneumonia severity on a large chest X-ray dataset
Alberto Signoroni, Mattia Savardi, Sergio Benini, Nicola Adami, Riccardo Leonardi, Paolo Gibellini, Filippo Vaccher, Marco Ravanelli, Andrea Borghesi, Roberto Maroldi, Davide Farina |
Medical Image Anal. | 5 |
| 2021 | Head pose estimation: A survey of the last ten years
Khalil Khan, Rehanullah Khan, Riccardo Leonardi, Pierangelo Migliorati, Sergio Benini |
Signal Process. Image Commun. | 3 |
| 2021 | Combining Appearance and Gradient Information for Image Symmetry DetectionabstractThis work addresses the challenging problem of reflection symmetry detection in unconstrained environments. Starting from the understanding on how the visual cortex manages planar symmetry detection, it is proposed to treat the problem in two stages: i) the design of a stable metric that extracts subsets of consistently oriented candidate segments, whenever the underlying 2D signal appearance exhibits definite near symmetric correspondences; ii) the ranking of such segments on the basis of the surrounding gradient orientation specularity, in order to reflect real symmetric object boundaries. Since these operations are related to the way the human brain performs planar symmetry detection, a better correspondence can be established between the outcomes of the proposed algorithm and a human-constructed ground truth. When compared to the testing sets used in recent symmetry detection competitions, a remarkable performance gain can be observed. In additional, further validation has been achieved by conducting perceptual validation experiments with users on a newly built dataset. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi |
IEEE Trans. Image Process. | 3 |
| 2021 | Learnable Descriptors for Visual SearchabstractThis work proposes LDVS, a learnable binary local descriptor devised for matching natural images within the MPEG CDVS framework. LDVS descriptors are learned so that they can be sign-quantized and compared using the Hamming distance. The underlying convolutional architecture enjoys a moderate parameters count for operations on mobile devices. Our experiments show that LDVS descriptors perform favorably over comparable learned binary descriptors at patch matching on two different datasets. A complete pair-wise image matching pipeline is then designed around LDVS descriptors, integrating them in the reference CDVS evaluation framework. Experiments show that LDVS descriptors outperform the compressed CDVS SIFT-like descriptors at pair-wise image matching over the challenging CDVS image dataset. Andrea Migliorati, Attilio Fiandrotti, Gianluca Francini, Riccardo Leonardi |
IEEE Trans. Image Process. | 4 |
| 2020 | 2D Discrete Mirror Transform for Image Non-Linear ApproximationabstractIn this paper, a new 2D transform named Discrete Mirror Transform (DMT) is presented. The DMT is computed by decomposing a signal into its even and odd parts around an optimal location in a given direction so that the signal energy is maximally split between the two components. After minimizing the information required to regenerate the original signal by removing redundant structures, the process is iterated leading the signal energy to distribute into a continuously smaller set of coefficients. The DMT can be displayed as a binary tree, where each node represents the single (even or odd) signal derived from the decomposition in the previous level. An optimized version of the DMT (ODMT) is also introduced, by exploiting the possibility to choose different directions at which performing the decomposition. Experimental simulations have been carried out in order to test the sparsity properties of the DMT and ODMT when applied on images: referring to both transforms, the results show a superior performance with respect to the popular Discrete Cosine Transform (DCT) and Discrete Wavelet Transform (DWT) in terms of non-linear approximation. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi |
ICPR | 3 |
| 2020 | Minimal Information Exchange for Secure Image Hash-Based Geometric Transformations EstimationabstractSignal processing applications dealing with secure transmission are enjoying increasing attention lately. This paper provides some theoretical insights as well as a practical solution for transmitting a hash of an image to a central server to be compared with a reference image. The proposed solution employs a rigid image registration technique viewed in a distributed source coding perspective. In essence, it embodies a phase encoding framework to let the decoder estimate the transformation parameters using a very modest amount of information about the original image. The problem is first cast in an ideal setting and then it is solved in a realistic scenario, giving more prominence to low computational complexity in both the transmitter and receiver, minimal hash size, and hash security. Satisfactory experimental results are reported on a standard images set. Fabrizio Guerrini, Marco Dalai, Riccardo Leonardi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Iterative Mirror Decomposition for Signal RepresentationabstractIn this paper it is shown how to describe any finite-energy continuous or discrete signal through an ordered set of positions to uniquely represent it. This is obtained by designing an iterative decomposition through a series of mirror operations around those positions. The purpose is to find at any step of the decomposition the location that provides for the maximum decoupling between the even and odd components of the signal with respect to it. The algorithm can then be iterated at infinity determining a sequence of positions. The per location information determines the optimal energy decoupling strategy at each stage providing remarkable sparsity in the representation. Thanks to the sparsity of the resulting representation, experimental simulations demonstrate superior approximation capabilities of this proposed non-linear mirror transform. Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi |
ICASSP | 3 |
| 2019 | Coding of Image Intra Prediction Residuals Using Symmetric GraphsabstractThe Discrete Cosine Transform (DCT) is widely deployed by modern image and video coding standards such as JPEG and H.26x. In most cases, the DCT is applied in a separable manner to rows and columns, which limits its ability to represent signals with diagonal orientation. As an alternative, non-separable transforms can represent signals with different orientations, but are significantly more computationally complex. To address this problem, in this paper we propose a set of non-separable Symmetry-Based Graph Fourier Transforms (SBGFTs), whose symmetric structures lead to a faster implementation. We study a practical image coding scenario that exploits the proposed SBGFTs, where for each intra predicted image residual block the optimal graph is chosen by solving a graph-based Rate-Distortion (R-D) problem. Experimental results indicate a coding efficiency higher than JPEG and JPEG2000. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 3 |
| 2019 | Face analysis through semantic face segmentation
Sergio Benini, Khalil Khan, Riccardo Leonardi, Massimo Mauro, Pierangelo Migliorati |
Signal Process. Image Commun. | 3 |
| 2018 | Symmetry-Based Graph Fourier Transforms for Image RepresentationabstractIt is well-known that the application of the Discrete Cosine Transform (DCT) in transform coding schemes is justified by the fact that it belongs to a family of transforms asymptotically equivalent to the Karhunen-Loeve Transform (KLT) of a first order Markov process. However, when the pixel-to-pixel correlation is low the DCT does not provide a compression performance comparable with the KLT. In this paper, we propose a set of symmetry-based Graph Fourier Transforms (GFT) whose associated graphs present a totally or partially symmetric grid. We show that this family of transforms well represents both natural images and residual signals outperforming the DCT in terms of energy compaction. We also investigate how to reduce the cardinality of the set of transforms through an analysis that studies the relation between efficient symmetry-based GFTs and the directional modes used in H.265 standard. Experimental results indicate that coding efficiency is high. Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, Antonio Ortega |
ICIP | 3 |
| 2018 | Feature Fusion for Robust Patch Matching with Compact Binary DescriptorsabstractThis work addresses the problem of learning compact yet discriminative patch descriptors within a deep learning framework. We observe that features extracted by convolutional layers in the pixel domain are largely complementary to features extracted in a transformed domain. We propose a convolutional network framework for learning binary patch descriptors where pixel domain features are fused with features extracted from the transformed domain. In our framework, while convolutional and transformed features are distinctly extracted, they are fused and provided to a single classifier which thus jointly operates on convolutional and transformed features. We experiment at matching patches from three different dataset, showing that our feature fusion approach outperforms multiple state-of-the-art approaches in terms of accuracy, rate and complexity. Andrea Migliorati, Attilio Fiandrotti, Gianluca Francini, Skjalg Lepsøy, Riccardo Leonardi |
MMSP | 5 |
| 2018 | Hair detection, segmentation, and hairstyle classification in the wild
Umar Riaz Muhammad, Michele Svanera, Riccardo Leonardi, Sergio Benini |
Image Vis. Comput. | 3 |
| 2017 | Head pose estimation through multi-class face segmentationabstractThe aim of this work is to explore the usefulness of face semantic segmentation for head pose estimation. We implement a multi-class face segmentation algorithm and we train a model for each considered pose. Given a new test image, the probabilities associated to face parts by the different models are used as the only information for estimating the head orientation. A simple algorithm is proposed to exploit such probabilites in order to predict the pose. The proposed scheme achieves competitive results when compared to most recent methods, according to mean absolute error and accuracy metrics. Moreover, we release and make publicly available a face segmentation dataset1consisting of 294 images belonging to 13 different poses, manually labeled into six semantic regions, which we used to train the segmentation models. Khalil Khan, Massimo Mauro, Pierangelo Migliorati, Riccardo Leonardi |
ICME | 4 |
| 2017 | Rate-Accuracy Optimization of Deep Convolutional Neural Network ModelsabstractRecently, deep learning has enjoyed a great deal of success for computer vision problems due to its capability to model highly complex tasks, such as image classification, object detection, face recognition, among many others. Although these neural networks are nowadays very powerful, there is a huge amount of parameters (i.e. the model) that need to be learned and require considerable storage space and bandwidth during transmission. This paper addresses the problems of storage and transmission of large deep learning models by proposing a compression solution that is independent of the model being trained as well as the data used for training. An efficient compression framework for the parameters of a neural network, more precisely the weights that interconnect the different neurons, which consume a significant amount of resources (memory, storage and bandwidth) is proposed. Several quantization strategies are considered as well as a statistical models for the different layers of a neural network, which are exploited by an arithmetic coding engine. Experimental results show that up to 92% bitrate savings can be obtained with minimal impact in terms of image classification accuracy. Alessandro Filini, João Ascenso, Riccardo Leonardi |
ISM | 3 |
| 2017 | Even/odd decomposition made sparse: A fingerprint to hidden patterns
Fabrizio Guerrini, Alessandro Gnutti, Riccardo Leonardi |
Signal Process. | 3 |
| 2017 | Interactive Film RecombinationabstractIn this article, we discuss an innovative media entertainment application called Interactive Movietelling. As an offspring of Interactive Storytelling applied to movies, we propose to integrate narrative generation through artificial intelligence (AI) planning with video processing and modeling to construct filmic variants starting from the baseline content. The integration is possible thanks to content description using semantic attributes pertaining to intermediate-level concepts shared between video processing and planning levels. The output is a recombination of segments taken from the input movie performed so as to convey an alternative plot. User tests on the prototype proved how promising Interactive Movietelling might be, even if it was designed at a proof of concept level. Possible improvements that are suggested here lead to many challenging research issues. Fabrizio Guerrini, Nicola Adami, Sergio Benini, Alberto Piacenza, Julie Porteous, Marc Cavazza, Riccardo Leonardi |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2016 | Figaro, hair detection and segmentation in the wildabstractHair is one of the elements that mostly characterize people appearance. Being able to detect hair in images can be useful in many applications, such as face recognition, gender classification, and video surveillance. To this purpose we propose a novel multi-class image database for hair detection in the wild, called Figaro. We tackle the problem of hair detection without relying on a-priori information related to head shape and location. Without using any human-body part classifier, we first classify image patches into hair vs. non-hair by relying on Histogram of Gradients (HOG) and Linear Ternary Pattern (LTP) texture features in a random forest scheme. Then we obtain results at pixel level by refining classified patches by a graph-based multiple segmentation method. Achieved segmentation accuracy (85%) is comparable to state-of-the-art on less challenging databases. Michele Svanera, Umar Riaz Muhammad, Riccardo Leonardi, Sergio Benini |
ICIP | 3 |
| 2016 | Shot scale distribution in art films
Sergio Benini, Michele Svanera, Nicola Adami, Riccardo Leonardi, András Bálint Kovács |
Multim. Tools Appl. | 4 |
| 2015 | Adaptive quantisation in HEVC for contouring artefacts removal in UHD contentabstractContouring artefacts affect the visual experience of some particular types of compressed Ultra High Definition (UHD) sequences characterised by smoothly textured areas and gradual transitions in the value of the pixels. This paper proposes a technique to adjust the quantisation process at the encoder so that contouring artefacts are avoided. The devised method does not require any change at the decoder side and introduces a negligible coding rate increment (up to 3.4% for the same objective quality). This result compares favourably with the average 11.2% bit-rate penalty introduced by a method where the quantisation step is reduced in contour-prone areas. Nicolo Casali, Matteo Naccari, Marta Mrak, Riccardo Leonardi |
ICIP | 4 |
| 2015 | Multi-class semantic segmentation of facesabstractIn this paper the problem of multi-class face segmentation is introduced. Differently from previous works which only consider few classes - typically skin and hair - the label set is extended here to six categories: skin, hair, eyes, nose, mouth and background. A dataset with 70 images taken from MIT-CBCL and FEI face databases is manually annotated and made publicly available1. Three kind of local features - accounting for color, shape and location - are extracted from uniformly sampled square patches. A discriminative model is built with random decision forests and used for classification. Many different combinations of features and parameters are explored to find the best possible model configuration. Our analysis shows that very good performance (~ 93% in accuracy) can be achieved with a fairly simple model. Khalil Khan, Massimo Mauro, Riccardo Leonardi |
ICIP | 3 |
| 2015 | Harmonic Change Detection for musical chords segmentationabstractIn this paper, different strategies for the calculation of the Harte's Harmonic Change Detection Function (HCDF) are discussed. HCDFs can be used for detecting chord boundaries for Automatic Chord Estimation (ACE) tasks, where the chord transitions are identified as peaks in the HCDF. We show that different audio features and different novelty metric have significant impact on the overall accuracy results of a chord segmentation algorithm. Furthermore, we show that certain combination of audio features and novelty measures provide a significant improvement with respect to the current chord segmentation algorithms. Alessio Degani, Marco Dalai, Riccardo Leonardi, Pierangelo Migliorati |
ICME | 3 |
| 2015 | Comparison of tuning frequency estimation methods
Alessio Degani, Marco Dalai, Riccardo Leonardi, Pierangelo Migliorati |
Multim. Tools Appl. | 3 |
| 2014 | An Integer Linear Programming Model for View Selection on Overlapping Camera ClustersabstractMulti-View Stereo (MVS) algorithms scale poorly on large image sets, and quickly become unfeasible to run on a single machine with limited memory. Typical solutions to lower the complexity include reducing the redundancy of the image set (view selection), and dividing the image set in groups to be processed independently (view clustering). A novel formulation for view selection is proposed here. We express the problem with an Integer Linear Programming (ILP) model, where cameras are modeled with binary variables, while the linear constraints enforce the completeness of the 3D reconstruction. The solution of the ILP leads to an optimal subset of selected cameras. As a second contribution, we integrate ILP camera selection with a view clustering approach which exploits Leveraged Affinity Propagation (LAP). LAP clustering can efficiently deal with large camera sets. We adapt the original algorithm so that it provides a set of overlapping clusters where the minimum and maximum sizes and the number of overlapping cameras can be specified. Evaluations on four different dataset show our solution provides significant complexity reductions and guarantees near-perfect coverage, making large reconstructions feasible even on a single machine. Massimo Mauro, Hayko Riemenschneider, Alberto Signoroni, Riccardo Leonardi, Luc Van Gool |
3DV | 4 |
| 2014 | A unified framework for content-aware view selection and planning through view importance
Massimo Mauro, Hayko Riemenschneider, Alberto Signoroni, Riccardo Leonardi, Luc Van Gool |
BMVC | 4 |
| 2014 | A Pitch Salience Function Derived from Harmonic Frequency Deviations for Polyphonic Music Analysis
Alessio Degani, Riccardo Leonardi, Pierangelo Migliorati, Geoffroy Peeters |
DAFx | 2 |
| 2014 | SubPatch: random kd-tree on a sub-sampled patch set for nearest neighbor field estimationabstractWe propose a new method to compute the approximate nearest-neighbors field (ANNF) between image pairs using random kd-tree and patch set sub-sampling. By exploiting image coherence we demonstrate that it is possible to reduce the number of patches on which we compute the ANNF, while maintaining high overall accuracy on the final result. Information on missing patches is then recovered by interpolation and propagation of good matches. The introduction of the sub-sampling factor on patch sets also allows for setting the desired trade off between accuracy and speed, providing a flexibility that lacks in state-of-the-art methods. Tests conducted on a public database prove that our algorithm achieves superior performance with respect to PatchMatch (PM) and Coherence Sensitivity Hashing (CSH) algorithms in a comparable computational time. Fabrizio Pedersoli, Sergio Benini, Nicola Adami, Masahiro Okuda, Riccardo Leonardi |
ICMV | 5 |
| 2014 | XKin: an open source framework for hand pose and gesture recognition using kinect
Fabrizio Pedersoli, Sergio Benini, Nicola Adami, Riccardo Leonardi |
Vis. Comput. | 4 |
| 2013 | Overlapping camera clustering through dominant sets for scalable 3D reconstructionabstractIn this work we present a method for clustering large unordered sets of cameras. Our method uses camera view information available from Structure-from-Motion (SfM) for computing a set of overlapping clusters suited for Multi-View Stereo (MVS) reconstruction. Our formulation of the problem uses the game theoretic model of dominant sets to find competing clustering solutions with computational simplicity. The overlapping solutions ensure more robust partial reconstructions. Experimental evaluations show that our method produces more regular cluster and overlap configurations with respect to the state of the art. This allows more scalable and higher quality reconstructions, while speeding up 6 times with respect to a MVS which uses all images at once. c 2013. Massimo Mauro, Hayko Riemenschneider, Luc Van Gool, Riccardo Leonardi |
BMVC | 4 |
| 2013 | Tracking characters in movies within logical story unitsabstractIn this paper, we propose a methodology to allow movie character recognition and tracking within movie scenes. In detail, we present a combination of a tracking algorithm robust against the problems of the currently available face detection algorithms and a face recognition process. We test how effective the system is in terms of both face tracking effectiveness and precision-recall results obtained for the recognition of the main characters present in an input movie. Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi |
MMSP | 4 |
| 2013 | Classifying cinematographic shot types
Luca Canini, Sergio Benini, Riccardo Leonardi |
Multim. Tools Appl. | 3 |
| 2013 | Affective Recommendation of Movies Based on Selected Connotative FeaturesabstractThe apparent difficulty in assessing emotions elicited by movies and the undeniable high variability in subjects' emotional responses to film content have been recently tackled by exploring film connotative properties: the set of shooting and editing conventions that help in transmitting meaning to the audience. Connotation provides an intermediate representation that exploits the objectivity of audiovisual descriptors to predict the subjective emotional reaction of single users. This is done without the need of registering users' physiological signals. It is not done by employing other people's highly variable emotional rates, but by relying on the intersubjectivity of connotative concepts and on the knowledge of user's reactions to similar stimuli. This paper extends previous work by extracting audiovisual and film grammar descriptors and, driven by users' rates on connotative properties, creates a shared framework where movie scenes are placed, compared, and recommended according to connotation. We evaluate the potential of the proposed system by asking users to assess the ability of connotation in suggesting film content able to target their affective requests. Luca Canini, Sergio Benini, Riccardo Leonardi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | A comparison of state-of-the-art technologies for irreversible compression of large medical datasetsabstractIn this work we compare state-of-the-art 3D coding technologies which are suitable for compression of 3D (volumetric) medical imagery. Our aim is to assess the usage of these methods on representative large datasets, also keeping into account relevant features related to modern application requirements. This work wants to be a useful reference for those interested in medical image coding technologies evaluation and, more in general, in the interdisciplinary debate about the usage of irreversible compression technologies for an improved and cost effective handling of diagnostic imaging processes and infrastructures (PACS, teleradiology) involving large datasets. Reproducibility and possible extension of this study are guaranteed by the use of publicly available reference software and datasets. Alberto Signoroni, Mario Pezzoni, Claudia Tonoli, Riccardo Leonardi |
CBMS | 4 |
| 2012 | Overlay optimization for Peer-to-Peer scalable video streamingabstractVideo streaming with Peer-to-Peer (P2P) architectures and Scalable Video Coding (SVC) appears to be an interesting solution for an efficient streaming in the heterogeneous scenario of Internet applications. A key issue in such approach is the optimization of the bandwidth capacity of the P2P system. In this paper we propose an innovative approach for the network overlay optimization based on Integer Linear Programming. The proposed approach is particularly suitable in case of push-based solutions for video streaming using SVC with prioritized content. The simulation results show the flexibility of the proposed model when used to generate an overlay following different constraints and operational requirements. The usability and the computational complexity of the proposed method is also analyzed when the overlay includes an high number of peers. Livio Lima, Marco Dalai, Pierangelo Migliorati, Riccardo Leonardi |
ICIP | 4 |
| 2012 | XKin -: eXtendable hand pose and gesture recognition library for kinectabstractIn this work we provide an open-source framework for Kinect enabling more natural and intuitive hand-gesture communication between human and computer devices. The software package is endowed with useful tools for training the system to work with user-defined postures and gestures. The XKin project is fully implemented in C and freely available at https://github.com/fpeder/XKin under FreeBSD License. Our goal is to encourage contributions from other researchers and developers in building an open and effective system for empowering a natural modality for human-machine interaction. Fabrizio Pedersoli, Nicola Adami, Sergio Benini, Riccardo Leonardi |
ACM Multimedia | 4 |
| 2011 | A robust pipeline for rapid feature-based pre-alignment of dense range scansabstractAiming at reaching an interactive and simplified usage of high-resolution 3D acquisition systems, this paper presents a fast and automated technique for pre-alignment of dense range images. Starting from a multi-scale feature point extraction and description, a processing chain composed by feature matching and correspondence searching, ranking grouping and skimming is performed to select the most reliable correspondences over which the correct alignment is estimated. Pre-alignment is obtained in few seconds per million point images on a off-the-shelf PC architecture. The experimental setup aimed to demonstrate the system behavior with respect to a set of concurrent requirements and the obtained performance are significant in the perspective of a fast, robust and unconstrained 3D object reconstruction. Francesco Bonarrigo, Alberto Signoroni, Riccardo Leonardi |
ICCV | 3 |
| 2011 | Image coding with face descriptors embeddingabstractContent descriptors, useful for browsing and retrieval tasks, are generally extracted and treated as a separate entity with respect to the nature of the content itself. At the same time, conventional coding processes do not take into account information carried out by content descriptors. Content descriptors are closely related to the content itself, and they potentially can be used to exploit redundancy in entropy coding processes. Embedding content descriptors in the bitstream can reduce content description extraction load, and at the same time, it can reduce the rate associated to the compressed content and its description. In this paper an effective implementation of this approach is presented, where image descriptors are actively used in the coding process for exploiting redundancy. First of all, image areas containing faces are detected and encoded using a scalable method, where the base layer is represented by the corresponding eigenfaces, and the enhancement layer is formed by the prediction error. The remaining areas are then encoded by using a traditional approach. Simulations show that achievable compression performances are comparable with those provided by conventional, making the proposed approach very convenient for source coding and content description. Alberto Boschetti, Nicola Adami, Riccardo Leonardi, Masahiro Okuda |
ICIP | 3 |
| 2011 | Optimal rate adaptation with Integer Linear Programming in the scalable extension of H.264/AVCabstractAdaptation for scalable video is one of the recent challenges in video distribution over modern networks, which are heterogeneous both in terms of available bandwidth and user end terminal capability. Scalable Video Coding offers the possibility to adapt the content following the “quality layer” abstraction. In this work we present a new method to optimally define quality layers using Integer Linear Programming and distortion models. The performances of the proposed approach are comparable with the state-of-the-art methods, but they are obtained with strong complexity reduction and augmented flexibility. Livio Lima, Massimo Mauro, Tea Anselmo, Daniele Alfonso, Riccardo Leonardi |
ICIP | 5 |
| 2011 | 3D-PMDC: A parallelized morphological wavelet codec for 3D medical datasets and teleradiology applicationsabstractModern medical imaging modalities produce increasingly large datasets. This trend can be in contrast with the computation and transmission time requirements coming from critical teleradiology applications. Multidimensional image compression techniques can be considered as enabling solutions on condition that they are able to guarantee a suitable combination of rate-distortion and computational performance which fulfill all the application domain requirements. In this work, we present a parallel version of our 3D Embedded Morphological Dilation Coding algorithm that allows a significant reduction of computation costs and the concurrent conservation of coding performance and of other relevant bitstream properties. A comparison with the recently released JPEG2000 part 10 (JP3D) standard put in evidence the value of the proposed solution, especially for teleradiology applications over heterogeneous networks. Alberto Signoroni, Mario Pezzoni, Riccardo Leonardi |
ICIP | 3 |
| 2011 | Generating story variants with constrained video recombinationabstractWe present a novel approach to the automatic generation of filmic variants within an implemented Video-Based Storytelling (VBS) system that successfully integrates video segmentation with stochastically controlled re-ordering techniques and narrative generation via AI planning. We have introduced flexibility into the video recombination process by sequencing video shots in a way that maintains local video consistency and this is combined with exploitation of shot polysemy to enable shot reuse in a range of valid semantic contexts. Results of evaluations on output narratives using a shared set of video data show consistency in terms of local video sequences and global causality with no loss of generative power. Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi, Julie Porteous, Jonathan Teutenberg, Marc Cavazza |
ACM Multimedia | 4 |
| 2011 | Changing video arrangement for constructing alternative storiesabstractCurrently, automatic generation of filmic variants faces a number of key technical issues and thus it usually resorts to the shooting of multiple versions of alternative scenes. However, recent advancements in video analysis has made this objective feasible, though semantic consistency must be somehow preserved. This demo presents a video-based storytelling (VBS) system that successfully integrates video processing with narrative generation by means of a shared semantic description. The novel filmic variants are constructed through a flexible video recombination process that takes advantage of the polysemy of baseline video segments. The short output video clips shown in this demo prove how the generated narratives are semantically consistent while keeping generative power intact. Alberto Piacenza, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi, Jonathan Teutenberg, Julie Porteous, Marc Cavazza |
ACM Multimedia | 4 |
| 2011 | The art of video MashUp: supporting creative users with an innovative and smart application
Daniela Cardillo, Amon Rapp, Sergio Benini, Luca Console, Rossana Simeoni, Elena Guercio, Riccardo Leonardi |
Multim. Tools Appl. | 7 |
| 2011 | High Dynamic Range Image Watermarking Robust Against Tone-Mapping OperatorsabstractHigh dynamic range (HDR) images represent the future format for digital images since they allow accurate rendering of a wider range of luminance values. However, today special types of preprocessing, collectively known as tone-mapping (TM) operators, are needed to adapt HDR images to currently existing displays. Tone-mapped images, although of reduced dynamic range, have nonetheless high quality and hence retain some commercial value. In this paper, we propose a solution to the problem of HDR image watermarking, e.g., for copyright embedding, that should survive TM. Therefore, the requirements imposed on the watermark encompass imperceptibility, a certain degree of security, and robustness to TM operators. The proposed watermarking system belongs to the blind, detectable category; it is based on the quantization index modulation (QIM) paradigm and employs higher order statistics as a feature. Experimental analysis shows positive results and demonstrates the system effectiveness with current state-of-art TM algorithms. Fabrizio Guerrini, Masahiro Okuda, Nicola Adami, Riccardo Leonardi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2011 | A Connotative Space for Supporting Movie Affective RecommendationabstractThe problem of relating media content to users' affective responses is here addressed. Previous work suggests that a direct mapping of audio-visual properties into emotion categories elicited by films is rather difficult, due to the high variability of individual reactions. To reduce the gap between the objective level of video features and the subjective sphere of emotions, we propose to shift the representation towards the connotative properties of movies, in a space inter-subjectively shared among users. Consequently, the connotative space allows to define, relate, and compare affective descriptions of film videos on equal footing. An extensive test involving a significant number of users watching famous movie scenes suggests that the connotative space can be related to affective categories of a single user. We apply this finding to reach high performance in meeting user's emotional preferences. Sergio Benini, Luca Canini, Riccardo Leonardi |
IEEE Trans. Multim. | 3 |
| 2010 | CBCD based on color features and landmark MDS-assisted distance estimationabstractContent-Based Copy Detection (CBCD) of digital videos is an important research field that aims at the identification of modified copies of an original clip, e.g., on the Internet. In this application, the video content is uniquely identified by the content itself, by extracting some compact features that are robust to a certain set of video transformations. Given the huge amount of data present in online video databases, the computational complexity of the feature extraction and comparison is a very important issue. In this paper, a landmark based multi-dimensional scaling technique is proposed to speed up the detection procedure which is based on exhaustive search and the MPEG-7 Dominant Color Descriptor. The method is evaluated under the MPEG Video Signature Core Experiment conditions, and simulation results show impressive time savings at the cost of a slightly reduced detection performance. Marzia Corvaglia, Fabrizio Guerrini, Riccardo Leonardi, Pierangelo Migliorati, Eliana Rossi |
ICASSP | 3 |
| 2010 | Flexible and effective High Dynamic Range image codingabstractThis paper presents an algorithm based on a two-layer coding scheme, where the original information is represented by means of a Low Dynamic Range (LDR) image, obtained by applying a tone mapping operator to the original HDR (High Dynamic Range), plus an enhancement layer, which allows to recover the full dynamic range. More specifically, the original HDR is represented with a format similar to the well known RGBE, which uses a shared exponent to compactly represent floating point numbers. With respect to the original RGBE, here the mantissa is composed by an approximation of the LDR image while the shared exponent represents the enhancement layer. This particular choice allows to split the original HDR into a color image, the mantissa, and a grayscale image, the exponent, which are very smooth signals that can be efficiently compressed by conventional image coding methods. With respect to already proposed similar schemas, two desirable features are then provided: a high coding efficiency, combined with the possibility to retrieve a displayable version of the original content. Alberto Boschetti, Nicola Adami, Riccardo Leonardi, Masahiro Okuda |
ICIP | 3 |
| 2010 | Toward a multi-feature approach to Content-Based Copy DetectionabstractVideo Content-Based Copy Detection (CBCD) is an emergent research field which is targeted to the identification of modified copies of an original clip in a given dataset, e.g., on the Internet. As opposed to digital watermarking, the content itself is used to uniquely identify the video through the extraction of features that need to be robust against a certain set of predetermined video attacks. This paper advocates the use of multiple features together with detection performance estimation to construct a flexible video signature instead of a fixed, single feature based one. To combine diverse features, a normalized linear combination is also proposed. The system performance boost is evaluated through the MPEG Video Signature Core Experiment dataset and experimental results show how the proposed signature scheme can achieve impressive improvements with respect to the single feature approach. Marzia Corvaglia, Fabrizio Guerrini, Riccardo Leonardi, Pierangelo Migliorati, Eliana Rossi |
ICIP | 3 |
| 2010 | Watershed segmentation of medical volumes with paint drop markingabstractWe present an improvement of the classical marker-controlled watershed approach in the direction of a better exploitation of user-defined markers. The combined action of a partial flooding and paint drops falling downwards on the gray value relief from marker locations, leads to a robust and meaningful identification of the candidate basins, which is a prerequisite for an accurate segmentation. This is useful for user-controlled segmentation of biomedical volumes in that it facilitates robust identification of complex 3D structures with inhomogeneous borders. To this end, a visual interactive segmentation system has been implemented where different user-data interaction tools can be selected by physicians to generate machine-understandable knowledge in a quick and compact way. Experimental results on selected use-cases demonstrate the strengths of the proposed solutions. Alberto Signoroni, G. Zanetti, R. Grazioli, Riccardo Leonardi |
ICIP | 4 |
| 2010 | Estimating cinematographic scene depth in movie shotsabstractIn film-making, the distance from the camera to the subject greatly affects the narrative power of a shot. By the alternate use of Long shots, Medium and Close-ups the director is able to provide emphasis on key passages of the filmed scene. In this work we investigate four different inherent characteristics of single shots which contain indirect information about scene depth, without the need to recover the 3D structure of the scene as a prior step. Specifically, 2D scene geometric composition, frame colour properties, shot motion distribution and content are considered for classifying shots into three main categories. In the experimental phase, using SVM classifiers, we first test the ability of each single feature in distinguishing shot types; then we combine the whole feature set in order to improve the classification performance. Sergio Benini, Luca Canini, Riccardo Leonardi |
ICME | 3 |
| 2010 | High Dynamic Range image tone mapping based on local Histogram EqualizationabstractHigh Dynamic Range (HDR) images can represent the acquired scene with a greater dynamic range of luminance than classical Low Dynamic Range (LDR) ones. Despite the recent diffusion of some HDR camera models, HDR displays are not yet in the market. For this reason HDR images need to be adapted in order to be properly rendered through conventional devices. This operation mainly consists in a dynamic range compression realized by applying a Tone Mapping Operator (TMO). In this work, a new tone map algorithm, derived from the Contrast Limited Adaptive Histogram Equalization (CLAHE) technique, is presented. With respect to the original CLAHE, in the proposed implementation an adaptive contrast limit and a new strategy for the determination of local tone mapping functions have been introduced. The comparison between the obtained LDR images, and those produced by applying State of the Art TMOs, evidences how the main characteristic of the proposed algorithm is the ability to equally enhance visibility in both dark and bright areas. This could be, for example, a key feature in video surveillance applications and automotive safety camera systems. Alberto Boschetti, Nicola Adami, Riccardo Leonardi, Masahiro Okuda |
ICME | 3 |
| 2010 | Interactive storytelling via video content recombinationabstractIn the paper we present a prototype of video-based storytelling that is able to generate multiple story variants from a baseline video. The video content for the system is generated by an adaptation of forefront video summarisation techniques that decompose the video into a number of Logical Story Units (LSU) representing sequences of contiguous and interconnected shots sharing a common semantic thread. Alternative storylines are generated using AI Planning techniques and these are used to direct the combination of elementary LSU for output. We report early results from experiments with the prototype in which the reordering of video shots on the basis of their high-level semantics produces trailers giving the illusion of different storylines. Julie Porteous, Sergio Benini, Luca Canini, Fred Charles, Marc Cavazza, Riccardo Leonardi |
ACM Multimedia | 6 |
| 2010 | Embedded indexing in scalable video coding
Nicola Adami, Alberto Boschetti, Riccardo Leonardi, Pierangelo Migliorati |
Multim. Tools Appl. | 3 |
| 2010 | RUSHES - an annotation and retrieval engine for multimedia semantic units
Oliver Schreer, Ingo Feldmann, Isabel Alonso Mediavilla, Pedro Concejero, Abdul Hamid Sadka, Mohammad Rafiq Swash, Sergio Benini, Riccardo Leonardi, Tijana Janjusevic, Ebroul Izquierdo |
Multim. Tools Appl. | 8 |
| 2009 | The IRIS Network of Excellence: Future Directions in Interactive Storytelling
Marc Cavazza, Ronan Champagnat, Riccardo Leonardi |
ICIDS | 3 |
| 2009 | Emotional identity of moviesabstractIn the field of multimedia analysis, attempts that lead to an emotional characterization of content have been proposed. In this work we aim at defining the emotional identity of a feature movie by positioning it into an emotional space, as if it was a piece of art. The multimedia content is mapped into a trajectory whose coordinates are connected to filming and cinematographic techniques used by directors to convey emotions. The trajectory evolution over time provides a strong characterization of the movie, by locating different movies into different regions of the emotional space. The ability of this tool in characterizing content has been tested by retrieving emotionally similar movies from a large database, using IMDb genre classification for the evaluation of results. Luca Canini, Sergio Benini, Pierangelo Migliorati, Riccardo Leonardi |
ICIP | 4 |
| 2008 | Fast dialogue indexing based on structure informationabstractIn this paper we propose a fast and simple method to index and search for dialogues in a video database. Dialogues are indexed relying only on the typical dialogue scene structure, which has been modeled using a statistical framework. The related entropy rate is chosen as a compact index for capturing the specificity of the scene structure. Specifically, entropy rate effectively represents the number of expressed visual concepts and the alternating pattern in which concepts appear inside dialogue scenes. As demonstrated on a large video data set, modern video management systems would benefit from integrating such structure information for effective content organization and dialogue search task. Sergio Benini, Pierangelo Migliorati, Riccardo Leonardi |
ICIP | 3 |
| 2008 | Scalable coding of image collections with embedded descriptorsabstractWith the increasing popularity of repositories of personal images, the problem of effective encoding and retrieval of similar image collections has become very important. In this paper we propose an efficient method for the joint scalable encoding of image-data and visual-descriptors, applied to collections of similar images. From the generated compressed bit stream, it is possible to extract and decode the visual information at different granularity levels, enabling the so called ldquomidstream content accessrdquo. The proposed approach is based on the appropriate combination of vector quantization (VQ) and JPEG2000 image coding. Specifically, the images are encoded at a first draft level using an optimal visual-codebook, while the residual errors are encoded using a JPEG2000 approach. In this way, the codebook of the VQ is freely available as an efficient visual descriptor of the considered image collection. This scalable representation supports fast browsing and retrieval of image collections providing also a coding efficiency comparable with those of standard image coding methods. Nicola Adami, Alberto Boschetti, Riccardo Leonardi, Pierangelo Migliorati |
MMSP | 3 |
| 2008 | JPIP proxy server for remote browsing of JPEG2000 imagesabstractThe JPEG2000 image compression standard offers scalability features in support of remote browsing applications. In particular Part 9 of the JPEG2000 standard defines a protocol called JPIP for interactivity with JPEG2000 code-streams and files. In client-server application based on JPIP, a client does not directly interact with the compressed file, but formulates requests using a simple syntax which identifies the current ldquoFocus Windowrdquo. In this kind of application particularly useful could be a proxy server, that potentially can improve the performance of the system through a better use of the network infrastructure. The aim of this work is to propose a proxy server with JPIP capabilities and shows the benefits that can be brought to remote browsing applications. Livio Lima, David S. Taubman, Riccardo Leonardi |
MMSP | 3 |
| 2008 | Distributed Video Coding: Selecting the most promising application scenarios
Fernando Pereira 0001, Christine Guillemot, Touradj Ebrahimi, Riccardo Leonardi, Sven Klomp |
Signal Process. Image Commun. | 5 |
| 2008 | On Unique DecodabilityabstractIn this paper, we propose a revisitation of the topic of unique decodability and of some fundamental theorems of lossless coding. It is widely believed that, for any discrete sourceX, every ldquouniquely decodablerdquo block code satisfiesE[l(X1,X2,...,Xn)]gesH(X1,X2,...,Xn) whereX1,X2,...,Xnare the firstnsymbols of the source,E[l(X1,X2,...,Xn)] is the expected length of the code for those symbols, andH(X1,X2,...,Xn) is their joint entropy. We show that, for certain sources with memory, the above inequality only holds when a limiting definition of ldquouniquely decodable coderdquo is considered. In particular, the above inequality is usually assumed to hold for any ldquopractical coderdquo due to a debatable application of McMillan's theorem to sources with memory. We thus propose a clarification of the topic, also providing an extended version of McMillan's theorem to be used for Markovian sources. Marco Dalai, Riccardo Leonardi |
IEEE Trans. Inf. Theory | 2 |
| 2007 | New fast search algorithm for base layer of H.264 scalable video coding extensionabstractIn this contribution, a fast search motion estimation algorithm for H.264/AVC SVC (scalable video coding) base layer with hierarchical B-frame structure for temporal decomposition is presented and compared with fast search motion estimation algorithm in JSVM software, that is the reference software for H.264/AVC SVC. The proposed technique is a block-matching based motion estimation algorithm working in two steps, called Coarse search and Fine search. The Coarse search is performed for each frame in display order, and for each 16x16 macroblock chooses the best motion vector at half pel accuracy. Fine search is performed for each frame in encoding order and finds the best prediction for each block type, reference frame and direction, choosing the best motion vector at quarter pel accuracy using R-D optimization. Both Coarse and Fine Search test 3 spatial and 3 temporal predictors, and add to the best one a set of updates. Livio Lima, Daniele Alfonso, Luca Pezzoni, Riccardo Leonardi |
DCC | 4 |
| 2007 | Distributed Coding of Shifts using the DFT PhaseabstractIn this paper we consider the problem of image encoding with side information at the decoder, where the side information is an integer shifted version of the image at the encoder. The encoder is asked to send the shift of its own image with respect to the side information which is only available at the decoder. We propose a solution based on the encoding of the phase sign of the DFT coefficients, taken at exponentially spaced positions. We first introduce the method under ideal hypothesis, i.e. noiseless conditions without border effects, giving a theoretical foundation to the technique. Then, we consider the more realistic case of noisy images with border effects, showing the effectiveness of the proposed method. Marco Dalai, Riccardo Leonardi, Pier Luigi Dragotti |
ICASSP (1) | 2 |
| 2007 | Wavelet-Based Encoding for HD ApplicationsabstractIn the past decades, most of the research on image and video compression has focused on addressing high bandwidth-constrained environments. However, for high resolution and high quality image and video compression, as in the case of high definition television (HDTV) or digital cinema (DC), the primary constraints are related to quality and flexibility. This paper presents a comparison between scalable wavelet-based video codecs and the state of the art in single point encoding and it investigates the obtainable compression efficiency when using temporal correlation with respect to pure intra coding. Livio Lima, Francesca Manerba, Nicola Adami, Alberto Signoroni, Riccardo Leonardi |
ICME | 5 |
| 2007 | State-of-the-Art and Trends in Scalable Video Compression With Wavelet-Based ApproachesabstractScalable video coding (SVC) differs form traditional single point approaches mainly because it allows to encode in a unique bit stream several working points corresponding to different quality, picture size and frame rate. This work describes the current state-of-the-art in SVC, focusing on wavelet based motion-compensated approaches (WSVC). It reviews individual components that have been designed to address the problem over the years and how such components are typically combined to achieve meaningful WSVC architectures. Coding schemes which mainly differ from the space-time order in which the wavelet transforms operate are here compared, discussing strengths and weaknesses of the resulting implementations. An evaluation of the achievable coding performances is provided considering the reference architectures studied and developed by ISO/MPEG in its exploration on WSVC. The paper also attempts to draw a list of major differences between wavelet based solutions and the SVC standard jointly targeted by ITU and ISO/MPEG. A major emphasis is devoted to a promising WSVC solution, named STP-tool, which presents architectural similarities with respect to the SVC standard. The paper ends drawing some evolution trends for WSVC systems and giving insights on video coding applications which could benefit by a wavelet based approach. Nicola Adami, Alberto Signoroni, Riccardo Leonardi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Improving Turbo Codec Integration in Pixel-Domain Distributed Video CodingabstractThe field of distributed video coding (DVC) theory has received a lot of attention in recent years and effective encoding techniques have been proposed. In the present work the framework of pixel domain Wyner-Ziv coding of video frames is considered, following the scheme proposed in A. Aaron et al. (2002). Some key frames are supposed to be available at the decoder while other frames are Wyner-Ziv encoded using turbo codes; at the decoder motion compensated interpolation between the key frames is performed in order to construct the side information for the Wyner-Ziv frame decoding. In this paper an improved model for the correlation noise between the side information frame and the original one is proposed. It is shown that modeling the nonstationary nature of the noise leads to substantial gain in the rate-distortion performance. Furthermore, by considering the memory of the noise, we show that some further gain can be obtained by placing an interleaver before the turbo codec so as to spread the correlation noise all over the frame Marco Dalai, Riccardo Leonardi, Fernando Pereira 0001 |
ICASSP (2) | 2 |
| 2006 | Image Watermarking Robust Against Non-Linear Value-Metric Scaling Based on Higher Order StatisticsabstractA new QIM-based image watermarking system for still images is proposed. The new system is expressly designed to cope with non-linear value-metric scaling attacks such as histogram stretching and gamma correction. By recognizing that any value-metric scaling attack must not change the global appearance of the image, we argue that the watermark should be inserted into high level visual features. We move a first step into this direction by proposing a system embedding the watermark into the kurtosis of selected image blocks. Though the kurtosis is not strictly invariant against non-linear gain, its value tends to remain constant whenever the image content is not altered significantly. The experiments we carried out confirm the validity of the new system, though some problems still need to be solved to make it suitable for real applications Fabrizio Guerrini, Riccardo Leonardi, Mauro Barni |
ICASSP (5) | 2 |
| 2006 | Extraction of Significant Video Summaries by Dendrogram AnalysisabstractIn the current video analysis scenario, effective clustering of shots facilitates the access to the content and helps in understanding the associated semantics. This paper introduces a cluster analysis on shots which employs dendrogram representation to produce hierarchical summaries of the video document. Vector quantization codebooks are used to represent the visual content and to group the shots with similar chromatic consistency. The evaluation of the cluster codebook distortions, and the exploitation of the dependency relationships on the dendrograms, allow to obtain only a few significant summaries of the whole video. Finally the user can navigate through summaries and decide which one best suites his/her needs for eventual post-processing. The effectiveness of the proposed method is demonstrated by testing it on a collection of video-data from different kinds of programmes. Results are evaluated in terms of metrics that measure the content representational value of the summarization technique. Sergio Benini, Aldo Bianchetti, Riccardo Leonardi, Pierangelo Migliorati |
ICIP | 3 |
| 2006 | Hierarchical Summarization of Videos by Tree-Structured Vector QuantizationabstractAccurate grouping of video shots could lead to semantic indexing of video segments for content analysis and retrieval. This paper introduces a novel cluster analysis which, depending both on the video genre and the specific user needs, produces a hierarchical representation of the video only on a reduced number of significant summaries. An outlook on a possible implementation strategy is then suggested. Specifically, vector-quantization codebooks are used to represent the visual content and to cluster the shots with a similar chromatic consistency. The evaluation of the codebook distortion introduced in each cluster is used to stop the procedure on few levels, exploiting the dependency relationships between clusters. Finally, the user can navigate through summaries at each hierarchical level and then decide which level to adopt for eventual post-processing. The effectiveness of the proposed method is validated through a series of experiments on real visual-data excerpted from different kinds of programmes Sergio Benini, Aldo Bianchetti, Riccardo Leonardi, Pierangelo Migliorati |
ICME | 3 |
| 2006 | Cyclostationary error analysis and filter properties in a 3D wavelet coding framework
Riccardo Leonardi, Alberto Signoroni |
Signal Process. Image Commun. | 1 |
| 2005 | The Future-Viewer visual environment for semantic characterization of video sequencesabstractIn this paper we address the problem of semi-automatic annotation of audio-visual sequences. Specifically, we propose the use of an innovative graphic framework, named Future-Viewer, to perform a quick and efficient annotation of a given multimedia document. The basic idea consists in visualising a 2-dimensional feature space in which the shots of the considered video sequence are located. In this window, shots with similar content fall near each other, and the proposed tool offers various functionalities for automatically and semi-automatically finding and annotating the shot clusters in such feature space. The proposed system has been used to analyze the content in terms of logical story units of few video sequences and the obtained results appear very interesting. Marco Campanella, Riccardo Leonardi, Pierangelo Migliorati |
ICIP (1) | 2 |
| 2005 | An Intuitive Graphic Environment for Navigation and Classification of Multimedia DocumentsabstractIn this work, we propose an intuitive graphic framework for the effective visualization of MPEG-7 low-level features, in the context of classification and annotation of audio-visual documents. This graphic tool is proposed to facilitate the access to the content, and to improve a quick understanding of the semantics associated to the considered document. The main visualization paradigm employed consists in representing a 2D feature space in which the shots of the audiovisual document are located. In another window, the same shots are drawn in a temporal bar that gives the users also the information related to the time domain. In the main window, shots with similar content fall near each other, and the proposed tool offers various functionalities for automatically and semi-automatically finding and annotating shot clusters in the feature space. The use of the proposed system to analyze the content of few video sequences has shown very interesting capabilities Marco Campanella, Riccardo Leonardi, Pierangelo Migliorati |
ICME | 2 |
| 2005 | Non prefix-free codes for constrained sequencesabstractIn this paper we consider the use of variable length non prefix-free codes for coding constrained sequences of symbols. We suppose to have a Markov source where some state transitions are impossible, i.e. the stochastic matrix associated with the Markov chain has some null entries. We show that classic Kraft inequality is not a necessary condition, in general, for unique decodability under the above hypothesis and we propose a relaxed necessary inequality condition. This allows, in some cases, the use of non prefix-free codes that can give very good performance, both in terms of compression and computational efficiency. Some considerations are made on the relation between the proposed approach and other existing coding paradigms Marco Dalai, Riccardo Leonardi |
ISIT | 2 |
| 2004 | Efficient (piecewise) linear minmax approximation of digital signalsabstractEfficient geometric algorithms are provided for the linear approximation of digital signals under the uniform norm. Given a set of n points (x/sub i/, y/sub i/), i=1..n, with x/sub i/<x/sub j/ if i Marco Dalai, Riccardo Leonardi |
ICASSP (2) | 2 |
| 2004 | l∞ norm based second generation image codingabstractMany second generation image coding techniques have been studied in recent years. Most of these methods consider the l/sub 2/ norm of the error introduced in the coded image, while for the l/sub /spl infin// case only predictive or transform based methods were considered up to now, focusing on near-lossless coding. In this paper we present a first scheme for l/sub /spl infin// norm in the framework of second generation image coding. The image is adaptively segmented into rectangular regions of varying size leading to a binary tree decomposition. The grey levels of the pixels within every leaf are approximated by means of l/sub /spl infin// sub-optimal bilinear surfaces. Marco Dalai, Riccardo Leonardi |
ICIP | 2 |
| 2004 | Inferring semantics from structural annotations of audio-visual documentsabstractIn this paper, a new approach for semantic extraction is proposed. Assuming that the semantics of interest associated to a multimedia document is subjective and that the user cannot easily construct a semantic description on different abstraction levels, we propose an interactive tool which allows to generate a semantic description by organizing an audio-visual document. The structural decomposition is the result of a guided annotation by the user: the user segments the input sequence in events, assigns each event to a specific class and includes other informations such as time, place and contained objects. The classification process can evolve dynamically, which means that the user can organize the semantics with various personalized and more specialized classes. Using the resulting structural descriptions and classification, our method automatically generates a richer semantic description. The system is totally MPEG-7 complaint. Nicola Adami, Marzia Corvaglia, Riccardo Leonardi |
MMSP | 3 |
| 2004 | Semantic indexing of soccer audio-visual sequences: a multimodal approach based on controlled Markov chainsabstractContent characterization of sport videos is a subject of great interest to researchers working on the analysis of multimedia documents. In this paper, we propose a semantic indexing algorithm which uses both audio and visual information for salient event detection in soccer. The video signal is processed first by extracting low-level visual descriptors directly from an MPEG-2 bit stream. It is assumed that any instance of an event of interest typically affects two consecutive shots and is characterized by a different temporal evolution of the visual descriptors in the two shots. This motivates the introduction of a controlled Markov chain to describe such evolution during an event of interest, with the control input modeling the occurrence of a shot transition. After adequately training different controlled Markov chain models, a list of video segments can be extracted to represent a specific event of interest using the maximum likelihood criterion. To reduce the presence of false alarms, low-level audio descriptors are processed to order the candidate video segments in the list so that those associated to the event of interest are likely to be found in the very first positions. We focus in particular on goal detection, which represents a key event in a soccer game, using camera motion information as a visual cue and the "loudness" as an audio descriptor. The experimental results show the effectiveness of the proposed multimodal approach. Riccardo Leonardi, Pierangelo Migliorati, Maria Prandini |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Semantic indexing of sports program sequences by audio-visual analysisabstractSemantic indexing of sports videos is a subject of great interest to researchers working on multimedia content characterization. Sports programs appeal to large audiences and their efficient distribution over various networks should contribute to widespread usage of multimedia services. In this paper, we propose a semantic indexing algorithm for soccer programs, which uses both audio and visual information for content characterization. The video signal is processed first by extracting low-level visual descriptors from the MPEG compressed bit-stream. The temporal evolution of these descriptors during a semantic event is supposed to be governed by a controlled Markov chain. This allows to determine a list of those video segments where a semantic event of interest is likely to be found, based on the maximum likelihood criterion. The audio information is then used to refine the results of the video classification procedure by ranking the candidate video segments in the list so that the segments associated to the event of interest appear in the very first positions of the ordered list. The proposed method is applied to goal detection. Experimental results show the effectiveness of the proposed cross-modal approach. Riccardo Leonardi, Pierangelo Migliorati, Maria Prandini |
ICIP (1) | 1 |
| 2003 | Overview of multimodal techniques for the characterization of sport programs
Nicola Adami, Riccardo Leonardi, Pierangelo Migliorati |
VCIP | 2 |
| 2003 | Selective coding with controlled quality decay for 2D and 3D images in a JPEG2000 framework
Alberto Signoroni, Fabio Lazzaroni, Riccardo Leonardi |
VCIP | 3 |
| 2003 | High-performance embedded morphological wavelet codingabstractIn this letter, an efficient morphological wavelet coder is proposed. The clustering trend of significant coefficients is captured by a new kind of multiresolution binary dilation operator. The layered and adaptive nature of the subband dilation makes it possible for the coding technique to produce an embedded bit-stream with a modest computational cost and state-of-the-art rate-distortion performance. Morphological wavelet coding appears promising because the localized analysis of wavelet coefficient clusters is adequate to capture intrinsic patterns of the source, which can have substantial benefits for reducing further the data redundancy. Fabio Lazzaroni, Riccardo Leonardi, Alberto Signoroni |
IEEE Signal Process. Lett. | 2 |
| 2002 | Semantics of multimedia in MPEG-7abstractIn this paper, we present the tools standardized by MPEG-7 for describing the semantics of multimedia. In particular, we focus on the abstraction model, entities, attributes and relations of MPEG-7 semantic descriptions. MPEG-7 tools can describe the semantics of specific instances of multimedia such as one image or one video segment but can also generalize these descriptions either to multiple instances of multimedia or to a set of semantic descriptions. The key components of MPEG-7 semantic descriptions are semantic entities such as objects and events, attributes of these entities such as labels and properties, and, finally, relations of these entities such as an object being the patient of an event. The descriptive power and usability of these tools has been demonstrated in numerous experiments and applications, these make them key candidates to enable intelligent applications that deal with multimedia at human levels. Ana B. Benitez, Hawley K. Rising, Corinne Jörgensen, Riccardo Leonardi, Alessandro Bugatti, Kôiti Hasida, Rajiv Mehrotra, A. Murat Tekalp, Ahmet Ekin, Toby Walker |
ICIP (1) | 4 |
| 2002 | Generating TV summaries for CE-devicesabstractAutomatically generated summaries of TV content are indispensable for content selection and navigation in CE-devices. We show two different types of summaries: Short video trailers and visual overviews consisting of representative frames. The demo does not only show the feasibility of the proposed algorithms, but also shows how different types of generated summaries can be used in future CE-devices. Gerhard Mekenkamp, Benoit Huet, Riccardo Leonardi |
ACM Multimedia | 3 |
| 2002 | Embedded morphological dilation coding for 2D and 3D images
Fabio Lazzaroni, Alberto Signoroni, Riccardo Leonardi |
VCIP | 3 |
| 2001 | Modeling and reduction of PSNR fluctuations in 3D wavelet codingabstractWe study the effects of the quantization of the three-dimensional wavelet transform coefficients, encoded using a 3D zerotree-based compression scheme. The processing of the quantization error operated by the synthesis filter bank actually determines a modulation of the reconstruction error statistics. This effect entails a non-negligible and potentially objectionable PSNR oscillation among adjacent slices of the decoded 3D data-set, which, e.g. in the biomedical field, would typically be used jointly to reach a diagnosis. In particular, we propose a cyclostationary model of the error statistics fluctuations and an experimental validation of the model itself. Moreover some strategies to reduce such oscillation are suggested and evaluated. Alberto Signoroni, Riccardo Leonardi |
ICIP (3) | 2 |
| 2001 | Evaluation Of Different Descriptors For Identifying Similar Video ShotsabstractIn this paper, three techniques for video sequence retrieval which use statistical measures of color patterns in shots are pro- posed and compared. The first technique is based on a correla- tion of the MPEG7 Dominant Color Descriptor (DC) [6] as single characteristic feature of a shot. The second approach is to model color pattern distribution in a shot with a codebook, obtained by VQ (Vector Quantization) of the frame blocks composing the shot. The last one models the color pattern distribution using a GMM (Gaussian Mixture Model). Such descriptors are used to establish correspondence between non consecutive camera records through an appropriately designed similarity measure. As such, a new dis- tance measure is used in the comparison between the shot descrip- tors, by extending the metric proposed in [9]. A comparison is made of the dissimilarity performance associated with each of the three proposed descriptors, demonstrating the superior results ob- tainable with the VQ based approach. Nicola Adami, Riccardo Leonardi, Yao Wang 0001 |
ICME | 2 |
| 2001 | Event Recognition In Sport Programs Using Low-Level Motion IndicesabstractIn this paper we present a semantic video indexing algo-rithm based on finite-state machines and low-level motion indices extracted from the MPEG compressed bit-stream. The problem of semantic video indexing is actually of great interest due to the wide diffusion of large video databases. In literature we can find many video indexing algorithms, based on various types of low-level features, but the prob-lem of semantic indexing is less studied and surely it is a great challenging one. The proposed algorithm is an exam-ple of solution to the problem of finding a semantic relevant event (e.g., scoring of a goal in a soccer game) in case of specific categories of audio-visual programmes. The sim-ulation results show that the proposed algorithm can effec-tively detect the presence of goals and other relevant events in sport programs. 1. A. Bonzanini, Riccardo Leonardi, Pierangelo Migliorati |
ICME | 2 |
| 2001 | Low level processing of audio and video information for extracting the semantics of contentabstractThe problem of semantic indexing of multimedia documents is actually of great interest due to the wide diffusion of large audio-video databases. We first briefly describe some techniques used to extract low-level features (e.g., shot change detection, dominant color extraction, audio classification etc.). Then the ToCAI (table of contents and analytical index) framework for content description of multimedia material is presented, together with an application which implements it. Finally we propose two algorithms suitable for extracting the high level semantics of a multimedia document. The first is based on finite-state machines and low-level motion indices, whereas the second uses hidden Markov models. Nicola Adami, Alessandro Bugatti, Riccardo Leonardi, Pierangelo Migliorati |
MMSP | 3 |
| 2001 | The ToCAI Description Scheme for Indexing and Retrieval of Multimedia Documents
Nicola Adami, Alessandro Bugatti, Riccardo Leonardi, Pierangelo Migliorati, Lorenzo A. Rossi |
Multim. Tools Appl. | 3 |
| 2000 | Describing multimedia documents in natural and semantic-driven ordered hierarchiesabstractIn this work we present the ToCAI (Table of Content-Analytical Index) framework, a description scheme (DS) for content description of audio-visual (AV) documents. The idea for such a description scheme comes out from the structures used for indexing technical books (table of content and analytical index). This description scheme provides therefore a hierarchical description of the time sequential structure of a multimedia document (ToC), suitable for browsing, together with an analytical index (AI) of the key items of the document, suitable for retrieval. The AI allows to represent in an ordered way the items of the AV document which are most relevant from the semantic point of view. The ordering criteria are therefore selected according to the application context. The detailed structure of the DS is presented by means of UML notation as well and an application example is shown. Nicola Adami, Alessandro Bugatti, Riccardo Leonardi, Pierangelo Migliorati, Lorenzo A. Rossi |
ICASSP | 3 |
| 2000 | Interactive Segmentation of Biomedical Images and Volumes Using Connected OperatorsabstractPresents a new method able to convey more meaningful images to the clinician. To reach this goal a new approach to medical interactive segmentation issue is proposed. First an automatic simplification process is performed on the volume, relying on connected operator methods. Afterwards a set of descriptors of the region of interest (RoI) is acquired by means of a 2D interactive session by the clinician. These descriptors, reflecting a real diagnostic interest, work like selective criteria during the subsequent morphological segmentation process done on the whole volume. They allow one to extract the desired structures avoiding a strong interaction by the user. Finally a 3D visualization of the segmented volume of interest (VoI) is obtained, and interactive tools allow the user to adjust and improve the results until satisfaction. Sergio Benini, E. Boniotti, Riccardo Leonardi, Alberto Signoroni |
ICIP | 3 |
| 1998 | Identification of Story Units in Audio-Visual Sequences by Joint Audio and Video ProcessingabstractA novel technique, which uses a joint audio-visual analysis for scene identification and characterization, is proposed. The paper defines four different scene types: dialogues, stories, actions, and generic scenes. It then explains how any audio-visual material can be decomposed into a series of scenes obeying the previous classification, by properly analyzing and then combining the underlying audio and visual information. A rule-based procedure is defined for such purpose. Before such rule-based decision can take place, a series of low-level pre-processing tasks are suggested to adequately measure audio and visual correlations. As far as visual information is concerned, it is proposed to measure the similarities between non-consecutive shots using a learning vector quantization approach. An outlook on a possible implementation strategy for the overall scene identification task is suggested, and validated through a series of experimental simulations on real audio-visual data. Caterina Saraceno, Riccardo Leonardi |
ICIP (1) | 2 |
| 1997 | Decimated wavelet representation of images-application to compressionabstractA new way to improve the representation of images using a discrete wavelet transform for coding purposes is presented. The idea lies in combining all wavelet coefficients related to detail information at a same resolution level but along different orientations (horizontal, vertical, and diagonal), into a single image. Given that detail information is located for all subband images in the neighborhood of high frequency textures or edge locations, the pattern of significant coefficients remains unchanged after the combination process. This process allows one to further reduce the number of transformed coefficients by 2/3, while preserving the multiresolution structure. This information can thus be efficiently coded using a multiresolution embedded coding scheme, such as Shapiro's (see IEEE Trans. on Signal Proc., vol.SP-41, no.12, p.3445-62, 1993) zerotree coder. Overall, a higher coding efficiency can be reached while preserving the cross-scale prediction of significance among the coefficients. Ultimately, approximate detail information must be recovered from the combined and coded data for each subband of the original wavelet, so as to reconstruct a decoded image. Riccardo Leonardi, Andrea Mazzarri, Alberto Signoroni |
ICASSP | 1 |
| 1997 | Audio as a support to scene change detection and characterization of video sequencesabstractA challenging problem to construct video databases is the organization of video information. The development of algorithms able to organize video information according to semantic content of the data is getting more and more important. This will allow algorithms such as indexing and retrieval to work more efficiently. Until now, an attempt to extract semantic information has been performed using only video information. As a video sequence is constructed from a 2-D projection of a 3-D scene, video processing has shown its limitations especially in solving problems such as object identification or object tracking, reducing the ability to extract semantic characteristics. A possibility to overcome the problem is to use additional information. The associated audio signal is then the most natural way to obtain this information. This paper presents a technique which combines video and audio information together for classification and indexing purposes. The classification is performed on the audio signal; a general framework that uses the results of such classification is then proposed for organizing video information. Caterina Saraceno, Riccardo Leonardi |
ICASSP | 2 |
| 1997 | Identification of Successive Correlated Camera Shots Using Audio and Video InformationabstractEffective creation and utilization of large digital multimedia libraries involves the application of multimedia processing techniques to organize, condense and index media for intelligent searching and selective retrieval. The segmentation of video sequences into scenes and the characterization of each scene has been suggested as a technique for organizing video information. As human beings use both auditory and visual systems to perceive the semantics of audio-visual sources, the analysis of the associated audio signal is shown to serve as a support for organizing video information. Caterina Saraceno, Riccardo Leonardi |
ICIP (3) | 2 |
| 1996 | Image compression using binary space partitioning treesabstractFor low bit-rate compression applications, segmentation-based coding methods provide, in general, high compression ratios when compared with traditional (e.g., transform and subband) coding approaches. In this paper, we present a new segmentation-based image coding method that divides the desired image using binary space partitioning (BSP). The BSP approach partitions the desired image recursively by arbitrarily oriented lines in a hierarchical manner. This recursive partitioning generates a binary tree, which is referred to as the BSP-tree representation of the desired image. The most critical aspect of the BSP-tree method is the criterion used to select the partitioning lines of the BSP tree representation, In previous works, we developed novel methods for selecting the BSP-tree lines, and showed that the BSP approach provides efficient segmentation of images. In this paper, we describe a hierarchical approach for coding the partitioning lines of the BSP-tree representation. We also show that the image signal within the different regions (resulting from the recursive partitioning) can be represented using low-order polynomials. Furthermore, we employ an optimum pruning algorithm to minimize the bit rate of the BSP tree representation (for a given budget constraint) while minimizing distortion. Simulation results and comparisons with other compression methods are also presented. Hayder Radha, Martin Vetterli, Riccardo Leonardi |
IEEE Trans. Image Process. | 3 |
| 1995 | Perceptual embedded image coding using wavelet transformsabstractWe present a modified version of an embedded wavelet coding scheme, first suggested by Shapiro (see IEEE Transactions on Signal Processing, vol.41, no. 12, p.3445-3462, 1993), that improves the performance of the original algorithm in a visual subjective distortion sense. We preserve the features of the original Shapiro's embedded coder. It is possible to choose a fixed target bit rate, as the information needed to represent an image coded at some rate always contains the needed information for the same image coded at lower rates. Therefore, the decoder can cease decoding the bit stream at any point, simulating an image coded at a lower rate corresponding to the truncated bit stream. We also introduce some perceptive improvements by adopting different (more regular) filters with respect to the original QMF pyramid filters proposed by Simoncelli, Hingorani et al. (1987) and used by Shapiro. These filters are synthesized using a "wavelet approach" instead of a "subband approach", and this leads to a better control on their regularity properties, jointly with better perceptual performance. Andrea Mazzarri, Riccardo Leonardi |
ICIP | 2 |
| 1994 | Tree Based Motion Compensated Video CodingabstractWe present a novel technique to encode video sequences, that performs a region-based decomposition of each frame on the basis of motion information. Using the segmentation map, any region in a frame to be encoded will be predicted from a single reference frame, using motion compensated prediction. The use of a single reference frame avoids feedback of the prediction error information in the prediction of successive frames. Coding is simply obtained by describing the segmentation map and the associated motion information. Error information will not be provided for low bit-rate applications. The segmentation map is described using a quadtree structure. Within such a tree structure, we show how motion information can be predicted either spatially or temporally, so as to minimize redundancy of information. The motion and segmentation information are estimated on the basis of a two stage process using the frame to be encoded and the reference frame: (1) a hierarchical top-down decomposition; (2) a bottom-up merging strategy. The proposed posed method is used to encode to encode QCIF video sequences with a reasonable duality at a 10 frame/s rate using roughly 20 kbit/s.> Riccardo Leonardi |
ICIP (2) | 1 |
| 1993 | Symmetrical segmentation-based image codingabstractAn image coding technique based on symmetry extraction and Binary Space Partitioning (BSP) tree representation for still pictures is presented. Axes of symmetry, detected through a principal axis of inertia approach and a coefficient of symmetry measure, are used to divide recursively an input image into a finite number of convex regions. This recursive partitioning results in the BSP tree representation of the image data. The iterative partition occurs whenever the current left/right node of the tree cannot be represented `symmetrically' by its counterpart, i.e., the right/left node. This splitting process may also end whenever the region associated with a given node has homogeneous characteristics or its size falls below a certain threshold. Given a BSP tree partition for a given input image, and the `seed' leaf nodes (i.e., those that cannot be generated by mirroring their counterparts), the remaining leaf nodes of the tree are reconstructed using a predictive scheme with respect to the `seed' leaf nodes. Christina Saraceno, Riccardo Leonardi |
VCIP | 2 |
| 1991 | A multiresolution approach to binary tree representations of imagesabstractA multiresolution method for constructing a BSP (binary space partitioning) tree is introduced. This approach derives a hierarchy (pyramid) of scale-space images from the original image. In this hierarchy, a BSP tree of an image is built from other trees representing low-resolution images of the pyramid. A low-resolution image BSP tree serves as an initial guess to construct a higher-resolution image tree. Due to filtering when constructing the pyramid, details are discarded. As a result, a more robust segmentation is obtained. Moreover a significant computational advantage is achieved.> Hayder Radha, Riccardo Leonardi, Martin Vetterli |
ICASSP | 2 |
| 1991 | Binary space partitioning tree representation of imagesabstractLTS Hayder Radha, Riccardo Leonardi, Martin Vetterli, Bruce Naylor |
J. Vis. Commun. Image Represent. | 2 |
| 1991 | Digital HDTV compression using parallel motion-compensated transform codersabstractThe authors suggest a parallel processing structure using the proposed international standard for visual telephony (CCITT P*64 kbs standard) as processing elements, to compress digital high definition television (HDTV) pictures. The basic idea is to partition an HDTV picture, in space or in frequency, into smaller sub-pictures and then compress each sub-picture using a CCITT P*64 kbs coder. This seems to be a cost-effective solution to the HDTV hardware. Since each sub-picture is processed by an independent coder, without coordination these coded sub-pictures may have unequal picture quality. To maintain a uniform quality HDTV picture, the following two issues are studied: sub-channel control strategy (bits allocated to each sub-picture); and quantization and buffer control strategy for individual sub-picture coders. Algorithms to resolve these problems and their computer simulations are presented.> Hsueh-Ming Hang, Riccardo Leonardi, Barry G. Haskell, Robert L. Schmidt, Hemant Bheda, Joseph H. Othmer |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1990 | Digital HDTV compression at 44 Mbps using parallel motion-compensated transform codersabstractHigh Definition Television (HDTV) promises to offer wide-screen, much better quality pictures as compared to the today’s television. However, without compression a digital HDTV channel may cost up to one Gbits/sec transmission bandwidth. We suggest a parallel processing structure using the proposed international standard for visual telephony (CCITT Px64 kbs standard) as processing elements, to compress the digital HDTV pictures. The basic idea is to partition an HDTV picture into smaller sub-pictures and then compress each sub-picture using a CCITT Px64kbs coder, which is cost-effective, by today’s technology, only on small size pictures. Since each sub-picture is processed by an independent coder, without coordination these coded sub-pictures may have unequal picture quality. To maintain a uniform quality HDTV picture, the following two issues are studied: (l) sub-channel control strategy (bits allocated to each sub-picture), and (2) quantization and buffer control strategy for individual sub-picture coder. Algorithms to resolve the above problems and their computer simulations are presented. Hsueh-Ming Hang, Riccardo Leonardi, Barry G. Haskell, Robert L. Schmidt, Hemant Bheda, Joseph H. Othmer |
VCIP | 2 |
| 1990 | Image representation using binary space partitioning treesabstractRepresentation of two and three-dimensional objects by tree structures has been used extensively in solid modeling, computer graphics, computer vision and image processing. (See for example [Mantyla] [Chen] [Hunter] [Rosenfeld] [Leonardi].) Quadtrees, which are used to represent objects in 2-D space, and octrees, which are the extension of quadtrees in 3-D space, have been studied thoroughly for applications in graphics and image processing. Hayder Radha, Riccardo Leonardi, Bruce Naylor, Martin Vetterli |
VCIP | 2 |
| 1990 | Video coding with motion-compensated interpolation for CD-ROM applications
Atul Puri, Rangarajan Aravind, Barry G. Haskell, Riccardo Leonardi |
Signal Process. Image Commun. | 4 |