Simone Milani

dblp:10/4730 · DBLP profile ↗
← Back
72ranked-venue papers
41as first author
13since 2021 · last 2026
0000-0001-8266-5839ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 38 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
abstract
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transition regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.
Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
IH&MMSec2
2026 TeLL-me what you cannot see: a vision-language framework for forensic mugshot augmentation
abstract
Abstract During criminal investigations, the availability of usable face imagery for persons of interest directly affects downstream investigative activities, including poster standardization and human-based search and review. In practice, agencies often face scarcity of high-quality images, heterogeneous capture conditions, and obsolescence, which can reduce the utility of available evidence and hinder timely information sharing. This paper introduces a forensic-oriented mugshot augmentation framework and evaluation protocol to support law-enforcement workflows with modern vision-language and generative models. The proposed modular pipeline optionally enhances low-quality inputs, extracts structured poster-style physical descriptors from a single image, and generates controlled synthetic portraits conditioned on those descriptors while monitoring identity consistency. By formalizing these steps and their associated measurements, the framework provides a reproducible reference for studying how such technologies behave under realistic constraints and for identifying failure cases relevant to forensic use. On the evaluated dataset, attribute extraction reached $$83.1\%$$ 83.1 % (+ 2.3 percentage points over the original mugshots), and re-identification backends showed clear separation between same-subject and different-subject pairs (similarity $$\sim$$ ∼ 0.70 vs. $$\sim$$ ∼ 0.50). These results indicate that the framework can improve descriptor faithfulness and support identity-consistent augmentation under the tested conditions; limitations and forensic-relevant risks (e.g., appearance drift and demographic sensitivity) are explicitly discussed.
Saverio Cavasin, Mattia Tamiazzo, Pietro Biasetton, Simone Milani, Mauro Conti
J. Inf. Secur.4
2025 Real or virtual: a video conferencing background manipulation-detection system
abstract
Abstract In the past few years, the popularity and wide use of video conferencing software enjoyed exponential growth in market size. This technology enables participants in different geographic regions to have a virtual face-to-face meeting. Additionally, it allows participants to utilize virtual backgrounds to hide their real environment with privacy concerns or to reduce distractions, particularly in professional settings. In scenarios where the users should not hide their actual locations, they may mislead other participants into assuming that the displayed virtual backgrounds are real. In this paper, we propose a new publicly-available dataset of virtual and real backgrounds in video conferencing software (e.g., Zoom, Google Meet, Microsoft Teams). The presented archive was evaluated by an exhaustive series of tests and scenarios using two well-known features extraction methods: CRSPAM1372 and six co-mat. The first verification scenario considers the case where the detector is unaware of manipulated frames (i.e., the forensically-edited frames are not part of the training set). A model trained on zoom frames that were tested with Google Meet frames can detect real background images from virtual ones in video conferencing software with 99.80% detection accuracy. Furthermore, it is possible to distinguish virtual from real backgrounds in videos created for videoconferencing software at a high detection rate of approximately 99.80%. According to our conclusions, the proposed method greatly enhanced the detection accuracy and resistance against diverse adversarial conditions, making it a reliable technique for classifying actual as opposed to virtual backgrounds in video communications. Given the described dataset provided and some preliminary experiments that we performed, we expect that it will lead to more future research in this domain.
Ehsan Nowroozi, Yassine Mekdad, Mauro Conti, Simone Milani, A. Selcuk Uluagac
Multim. Tools Appl.4
2025 Learning From Mistakes: Self-Regularizing Hierarchical Representations in Point Cloud Semantic Segmentation
abstract
Recent advances in autonomous robotic technologies have highlighted the growing need for precise environmental analysis. Point cloud semantic segmentation has gained attention to accomplish fine-grained scene understanding by acting directly on raw content provided by sensors. Recent solutions showed how different learning techniques can be used to improve the performance of the model, without any architectural or dataset change. Following this trend, we present a coarse-to-fine setup that LEArns from classification mistaKes (LEAK) derived from a standard model. First, classes are clustered into macro groups according to mutual prediction errors; then, the learning process is regularized by: (1) aligning class-conditional prototypical feature representation for both fine and coarse classes, (2) weighting instances with a per-class fairness index. Our LEAK approach is very general and can be seamlessly applied on top of any segmentation architecture; indeed, experimental results showed that it enables state-of-the-art performances on different architectures, datasets and tasks, while ensuring more balanced class-wise results and faster convergence.
Elena Camuffo, Umberto Michieli, Simone Milani
IEEE Trans. Multim.3
2024 Continual Road-Scene Semantic Segmentation Via Feature-Aligned Symmetric Multi-Modal Network
abstract
State-of-the-art multimodal semantic segmentation strategies combining LiDAR and color data are usually designed on top of asymmetric information-sharing schemes and assume that both modalities are always available. This strong assumption may not hold in real-world scenarios, where sensors are prone to failure or can face adverse conditions that make the acquired information unreliable. This problem is exacerbated when continual learning scenarios are considered since they have stringent data reliability constraints. In this work, we re-frame the task of multimodal semantic segmentation by enforcing a tightly coupled feature representation and a symmetric information-sharing scheme, which allows our approach to work even when one of the input modalities is missing. We also introduce an ad-hoc class-incremental continual learning scheme, proving our approach’s effectiveness and reliability even in safety-critical settings, such as autonomous driving. We evaluate our approach on the SemanticKITTI dataset, achieving impressive performances.
Francesco Barbato, Elena Camuffo, Simone Milani, Pietro Zanuttigh
ICIP3
2024 Point Cloud Geometry Scalable Coding with a Quality-Conditioned Latents Probability Estimator
abstract
The widespread usage of point clouds (PC) for immersive visual applications has resulted in the use of very heterogeneous receiving conditions and devices, notably in terms of network, hardware, and display capabilities. In this scenario, quality scalability, i.e., the ability to reconstruct a signal at different qualities by progressively decoding a single bitstream, is a major requirement that has yet to be conveniently addressed, notably in most learning-based PC coding solutions. This paper proposes a quality scalability scheme, named Scalable Quality Hyperprior (SQH), adaptable to learning-based static point cloud geometry codecs, which uses a Quality-conditioned Latents Probability Estimator (QuLPE) to decode a high-quality version of a PC learning-based representation, based on an available lower quality base layer. SQH is integrated in the future JPEG PC coding standard, allowing to create a layered bitstream that can be used to progressively decode the PC geometry with increasing quality and fidelity. Experimental results show that SQH offers the quality scalability feature with very limited or no compression performance penalty at all when compared with the corresponding non-scalable solution, thus preserving the significant compression gains over other state-of-the-art PC codecs.
Daniele Mari, André F. R. Guarda, Nuno M. M. Rodrigues, Simone Milani, Fernando Pereira 0001
ICIP4
2024 Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
abstract
Developing a reliable vision system is a fundamental challenge for robotic technologies (e.g., indoor service robots and outdoor autonomous robots) which can ensure reliable navigation even in challenging environments such as adverse weather conditions (e.g., fog, rain), poor lighting conditions (e.g., over/under exposure), or sensor degradation (e.g., blurring, noise), and can guarantee high performance in safety-critical functions. Current solutions proposed to improve model robustness usually rely on generic data augmentation techniques or employ costly test-time adaptation methods. In addition, most approaches focus on addressing a single vision task (typically, image recognition) utilising synthetic data. In this paper, we introduce Per-corruption Adaptation of Normalization statistics (PAN) to enhance the model robustness of vision systems. Our approach entails three key components: (i) a corruption type identification module, (ii) dynamic adjustment of normalization layer statistics based on identified corruption type, and (iii) real-time update of these statistics according to input data. PAN can integrate seamlessly with any convolutional model for enhanced accuracy in several robot vision tasks. In our experiments, PAN obtains robust performance improvement on challenging real-world corrupted image datasets (e.g., OpenLoris, ExDark, ACDC), where most of the current solutions tend to fail. Moreover, PAN outperforms the baseline models by 20-30% on synthetic benchmarks in object recognition tasks.
Elena Camuffo, Umberto Michieli, Simone Milani, Ji Joong Moon, Mete Ozay
IROS3
2024 Fingerprint membership and identity inference against generative adversarial networks
Saverio Cavasin, Daniele Mari, Simone Milani, Mauro Conti
Pattern Recognit. Lett.3
2023 Fully Automated Scan-to-BIM Via Point Cloud Instance Segmentation
abstract
Digital reconstruction through Building Information Models (BIM) is a valuable methodology for documenting and analyzing existing buildings. Its pipeline starts with geometric acquisition. (e.g., via photogrammetry or laser scanning) for accurate point cloud collection. However, the acquired data are noisy and unstructured, and the creation of a semantically-meaningful BIM representation requires a huge computational effort, as well as expensive and time-consuming human annotations. In this paper, we propose a fully automated scan-to-BIM pipeline. The approach relies on: (i) our dataset (HePIC), acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models; (ii) a novel ad hoc deep network (BIM-Net++) for semantic segmentation, whose output is then processed to extract instance information necessary to recreate BIM objects; (iii) novel model pre-training and class re-weighting to eliminate the need for a large amount of labeled data and human intervention.
D. Campagnolo, Elena Camuffo, Umberto Michieli, Paolo Borin, Simone Milani
ICIP5
2022 Hand Me Your PIN! Inferring ATM PINs of Users Typing with a Covered Hand
Matteo Cardaioli, Stefano Cecconello, Mauro Conti, Simone Milani, Stjepan Picek, Eugen Saraci
USENIX Security Symposium4
2021 Looking Through Walls: Inferring Scenes from Video-Surveillance Encrypted Traffic
abstract
Nowadays living environments are characterized by networks of inter-connected sensing devices that accomplish different tasks, e.g., video surveillance of an environment by a network of CCTV cameras. A malicious user could gather sensitive details on people’s activities by eavesdropping the exchanged data packets. To overcome this problem, video streams are protected by encryption systems, but even secured channels may still leak some information. In this paper, we show that it is possible to infer visual data by intercepting the encrypted video stream of a surveillance system, and how this may be leveraged to track the movements of a person inside the secured area. We trained an automatic classifier on a computer graphic simulator and tested it on real videos, with standard encryption protocols. Experiments proved the transferability of the classifier trained on synthetic sequences, succeeding in the detection of up to four different walking directions on real videos, with a limited amount of intercepted traffic.
Daniele Mari, Samuele Giuliano Piazzetta, Sara Bordin, Luca Pajola, Sebastiano Verde, Simone Milani, Mauro Conti
ICASSP6
2021 ADAE: Adversarial Distributed Source Autoencoder For Point Cloud Compression
abstract
The current paper presents an adversarial autoencoding strategy for voxelized point cloud geometry based on the principles of distributed source coding. The encoder characterizes the input voxel blocks with an array of hash bytes while the decoder combines them with side information blocks in order to reconstruct the original data. The reconstruction process is optimized by classifying the reconstructed block with an adversarial discriminator in order to make the recovered data as close as possible to an original block. Experimental results show that the proposed solution generalizes well while obtaining better coding performance with respect to other state-of-the-art solutions and allowing high flexibility in rate shaping and decoding operations.
Simone Milani
ICIP1
2021 A distributed source autoencoder of local visual descriptors for 3D reconstruction
Simone Milani
Pattern Recognit. Lett.1
2020 Phylogenetic Minimum Spanning Tree Reconstruction Using Autoencoders
abstract
The history of a shared and re-posted multimedia content can be reconstructed by analyzing the mutual relations between all of its near-duplicate copies and solving a minimum spanning tree (MST) problem, as shown by multimedia phylogeny research field, Unfortunately, MST estimation strategies are severely impaired by the noise affecting dissimilarity measures between pairs of near-duplicate contents, For this reason, researchers have recently been investigating robust dissimilarity metrics.This paper proposes a matrix denoising solution that both mitigates dissimilarity noise and reconstruct the desired phylogenetic tree at the same time, The proposed strategy is a first attempt to estimate a MST via a denoising autoencoder that returns an approximation of the adjacency matrix corresponding to the underlying tree, Experimental results prove that the proposed solution outperforms the previous approaches and easily adapts to different analysis scenarios.
Riccardo Castelletto, Simone Milani, Paolo Bestagini
ICASSP2
2020 A Syndrome-Based Autoencoder For Point Cloud Geometry Compression
abstract
Point cloud compression has been extensively-investigated in the past twenty years to find effective solutions that reduce the coded bit stream and permits adapting the coded bit rate to different scenarios. Despite these efforts, predictive strategies have so far performed poorly because of the low correlation level of the input data and the flexibility requirements, which imply minimizing the decoding dependences.The current paper proposes a convolutional autoencoder that applies the principles of Distributed Source Coding (DSC) to the deep representations of voxelized point cloud geometry data. The hidden variables, called syndromes, enable reconstructing the coded point cloud geometry from different reference data. The proposed strategy overcomes the state-of-the-art solutions in terms of flexibility and rate-distortion performance.
Simone Milani
ICIP1
2020 On the use of Benford's law to detect GAN-generated images
abstract
The advent of Generative Adversarial Network (GAN) architectures has given anyone the ability of generating incredibly realistic synthetic imagery. The malicious diffusion of GAN-generated images may lead to serious social and political consequences (e.g., fake news spreading, opinion formation, etc.). It is therefore important to regulate the widespread distribution of synthetic imagery by developing solutions able to detect them. In this paper, we study the possibility of using Benford's law to discriminate GAN-generated images from natural photographs. Benford's law describes the distribution of the most significant digit for quantized Discrete Cosine Transform (DCT) coefficients. Extending and generalizing this property, we show that it is possible to extract a compact feature vector from an image. This feature vector can be fed to an extremely simple classifier for GAN-generated image detection purpose.
Nicolò Bonettini, Paolo Bestagini, Simone Milani, Stefano Tubaro
ICPR3
2020 Ground-to-Aerial Viewpoint Localization via Landmark Graphs Matching
abstract
The capability of associating an image to its geographical location is a significant concern in journalism and digital forensics. Given the availability of geo-tagged satellite imagery for most of the Earth's surface, retrieving the location of a generic picture can be addressed as a cross-view image matching between aerial and ground views. In this paper, we outline some initial steps toward the development of a fully-unsupervised algorithm for ground-to-aerial image matching, exploiting the view-invariant adjacency relationships of the landmarks appearing in both views. We introduce a graph-based strategy that, given a set of pre-extracted landmarks, localizes the viewpoint of a ground-level 360-degree image within a broad aerial view of the same area, by matching the respective landmark graphs according to a specifically designed likelihood model.
Sebastiano Verde, Thiago Resek, Simone Milani, Anderson Rocha 0001
IEEE Signal Process. Lett.3
2020 A Transform Coding Strategy for Dynamic Point Clouds
abstract
The development of real-time 3D sensing devices and algorithms (e.g., multiview capturing systems, Time-of-Flight depth cameras, LIDAR sensors), as well as the widespreading of enhanced user applications processing 3D data, have motivated the investigation of innovative and effective coding strategies for 3D point clouds. Several compression algorithms, as well as some standardization efforts, has been proposed in order to achieve high compression ratios and flexibility at a reasonable computational cost. This paper presents a transform-based coding strategy for dynamic point clouds that combines a non-linear transform for geometric data with a linear transform for color data; both operations are region-adaptive in order to fit the characteristics of the input 3D data. Temporal redundancy is exploited both in the adaptation of the designed transform and in predicting the attributes at the current instant from the previous ones. Experimental results showed that the proposed solution obtained a significant bit rate reduction in lossless geometry coding and an improved rate-distortion performance in the lossy coding of color components with respect to state-of-the-art strategies.
Simone Milani, Enrico Polo, Simone Limuti
IEEE Trans. Image Process.1
2019 Material Identification Using RF Sensors and Convolutional Neural Networks
abstract
Recent years have assisted a widespreading of Radio-Frequency-based tracking and mapping algorithms for a wide range of applications, ranging from environment surveillance to human-computer interface. This work presents a material identification system based on a portable 3D imaging radar-based system, the Walabot sensor by Vayyar Technologies; the acquired three-dimensional radiance map of the analyzed object is processed by a Convolutional Neural Network in order to identify which material the object is made of. Experimental results show that processing the three-dimensional radiance volume proves to be more efficient thas processing the raw signals from antennas. Moreover, the proposed solution presents a higher accuracy with respect to some previous state-of-the-art solutions.
Gianluca Agresti, Simone Milani
ICASSP2
2019 Phylogenetic Analysis of Software Using Cache Miss Statistics
abstract
While the phylogenetic analysis of multimedia documents keeps being investigated, some recent studies have shown the possibility of re-using the same strategies to analyze the evolution of computer programs (Software Phylogeny), considering its many applications spanning from copyright enforcement to malware detection. This paper presents a solution for reconstructing the phylogenetic dependencies of different releases of a given program. The proposed method collects cache miss statistics during the program execution, builds a dissimilarity matrix from the results, and then estimates the corresponding Software Phylogenetic Tree (SPT) using a minimum spanning tree algorithm.
Sebastiano Verde, Simone Milani, Giancarlo Calvagno
ICASSP2
2019 Rendering-Aware Point Cloud Coding for Mixed Reality Devices
abstract
The recent diffusion of wearable and portable Augmented and Mixed Reality devices have highlighted some open problems in the transmission and visualization of three-dimensional point clouds such as the adaptation of the bit stream to different devices and networks or the optimization of the rendering/coding complexity.The current paper proposes a rendering-aware approach for the compression of static point cloud models that employs a multi-resolution representation of the model in spherical coordinates. This approach proves to be extremely effective in shaping the transmitted bit stream and rendering operations according to the complexity of the 3D model and the available calculation resources. Experimental results show that the proposed solution outperforms one of the most recent state-of-the-art solutions in terms of rate-distortion performance and computional effort.
Fabio Capraro, Simone Milani
ICIP2
2018 Improving Consensus-Based Distributed Camera Calibration Via Edge Pruning and Graph Traversal Initialization
abstract
Over the past few years, a huge number of distributed camera calibration strategies have been proposed for video surveillance and monitoring systems involving mobile terminals. Many of the proposed solutions rely on consensus-based algorithms, which aim at estimating the configuration of the network via a message passing protocol. In this paper we propose an improved consensus-based distributed camera calibration strategy that exploits a robust initialization, together with a pruning protocol to remove faulty links which could propagate excessively-noisy information through the network reducing the convergence time. The proposed solution seems to improve the state-of-the-art strategies in terms of accuracy, convergence speed, and computational complexity.
Giulia Michieletto, Simone Milani, Angelo Cenedese, Giacomo Baggio
ICASSP2
2018 A Transform Coding Strategy for Voxelized Dynamic Point Clouds
abstract
With the advent of virtual and augmented reality applications, 3D and free-viewpoint representations have evolved towards solid scene models using meshes and point clouds. Recent works have been addressing point clouds compression via octree-based hierarchical strategies in order to enable a multiresolution coding and visualization at a reasonable computational cost. This paper presents a voxelized dynamic point cloud coding scheme that combines a Cellular Automata block reversible transform for geometric data with a region adaptive transform for color data. Temporal redundancy is removed using a low-complexity prediction scheme to minimize the computational complexity and reduce the coded bit rate. Experimental results showed that the proposed solution obtained a significant bit rate reduction in lossless geometry coding and an improved rate-distortion performance in the lossy coding of color components with respect to state-of-the-art strategies.
Simone Limuti, Enrico Polo, Simone Milani
ICIP3
2018 Video Codec Forensics Based on Convolutional Neural Networks
abstract
The recent development of multimedia has made video editing accessible to everyone. Unfortunately, forensic analysis tools capable of detecting traces left by video processing operations in a blind fashion are still at their beginnings. One of the reasons is that videos are customary stored and distributed in a compressed format, and codec-related traces tends to mask previous processing operations. In this paper, we propose to capture video codec traces through convolutional neural networks (CNNs) and exploit them as an asset. Specifically, we train two CNN s to extract information about the used video codec and coding quality, respectively. Building upon these CNN s, we propose a system to detect and localize temporal splicing for video sequences generated from the concatenation of different video segments, which are characterized by inconsistent coding schemes and/or parameters (e.g., video compilations from different sources or broadcasting channels). The proposed solution is validated using videos at different resolutions (i.e., CIF, 4CIF, PAL and 720p) encoded with four common codecs (i.e., MPEG2, MPEG4, H264 and H265) at different qualities (i.e., different constant and variable bitrates, as well as constant quantization parameters).
Sebastiano Verde, Luca Bondi, Paolo Bestagini, Simone Milani, Giancarlo Calvagno, Stefano Tubaro
ICIP4
2017 3D reconstruction from web harvested images using a forensic quality metric
abstract
Structure-from-Motion (SfM) algorithms have recently been employed to reconstruct 3D scenes or environments from large sets of unordered images which were harvested from the web. Unfortunately, the accuracy of the reconstruction is significantly affected by the quality and the amount of editing operated on the processed images. Indeed, 3D modelling can significantly benefit from including forensic analysis strategies that are able to reconstruct the processing history of the processed images and select the most reliable pieces of visual information. The current paper presents an SfM strategy that orders the different views/images of the scene in the reconstruction process according to a processing age metric, i.e., a metric parameterizing the amount of processing stages operated on each image. Experimental results show that the proposed solution can improve the estimation accuracy of both 3D points and camera parameters with respect to state-of-the-art solutions.
Mattia Lecci, Simone Milani
ICASSP2
2017 Improving 3D reconstruction tracks using denoised euclidean distance matrices
abstract
The reconstruction of 3D point cloud models from unordered and uncalibrated sets of images has recently been a hot topic in the computer vision world. Most of the proposed solutions rely on the Structure-From-Motion algorithms, and their performances are significantly affected by the processing order (called track) of the considered images. This is computed according to a distance (or similarity) metric between couples of images, which is usually highly noisy. The paper proposes an image ordering strategy that models the distances between images as an Euclidean distance matrix and applies a rank-based denoising algorithm in order to refine the metric values. Experimental results prove that the accuracy of the final 3D model is sensibly improved.
Simone Milani
ICIP1
2017 Fast point cloud compression via reversible cellular automata block transform
abstract
Augmented and mixed reality applications require efficient tools permitting the compression and the visualization of 3D object at a limited computational cost. To this purpose, 3D point cloud representations have been widely used, together with an octree-based hierarchical organization of data that enables a multi-resolution visualization. This paper presents a voxel coding strategy based on a hierarchical Cellular Automata block reversible transform which permits obtaining a multi-resolution representation of the input volume and a higher compression gain with respect to the state-of-the-art octree strategies. The proposed solution also proves to be more flexible in defining multiple layers and more effective in preserving 3D volume quality when the stream is partially decoded.
Simone Milani
ICIP1
2016 Phylogenetic analysis of near-duplicate images using processing age metrics
abstract
Recent researches on image forensics have led to the design of algorithms to study the phylogenetic relationship between near-duplicate (ND) images. The proposed solutions aim at reconstructing the image phylogeny tree (IPT), and they have immediate applications in security, law and copyright enforcement, and news tracking services. Anyway, the effectiveness of such strategies strictly depends on the accuracy in characterizing image similarities. In this paper, we show that it is possible to take into account additional information to better reconstruct the IPT. More specifically, we propose a set of features that blindly model the processing age of an image, i.e., how much an image has been edited in its lifetime. By exploiting these features, it is possible to improve the performance of IPT reconstruction by increasing the accuracy and reducing the computational complexity.
Simone Milani, Marco Fontana, Paolo Bestagini, Stefano Tubaro
ICASSP1
2016 Depth map coding with elastic contours and 3D surface prediction
abstract
Depth maps are typically made of smooth regions separated by sharp edges. Following this rationale, this paper presents a novel coding scheme where depth data is represented by a set of contours defining the various regions together with a compact representation of the values inside each region. The proposed coding scheme is based on elastic curves, which make possible to compactly represent the contours exploiting also the temporal consistency in different frames. A 3D surface prediction algorithm is then used to obtain an accurate estimation of the depth field from the coded contours and a subsampled version of the data. Finally, an ad-hoc coding strategy for the low resolution data and the prediction residuals is presented. Experimental results prove how the proposed approach is able to obtain a very high coding efficiency outperforming the HEVC coder at medium-low bitrates.
Marco Calemme, Pietro Zanuttigh, Simone Milani, Marco Cagnazzo, Béatrice Pesquet-Popescu
ICIP3
2016 Impact of drone swarm formations in 3D scene reconstruction
abstract
Swarms of unmanned autonomous vehicles adopt effective algorithms to control their localization and formation. These strategies grant safe and collision-free navigation, control the positioning of drones, and automatize their motion according to the final task. This paper investigates how formations affect the reconstruction accuracy of a swarm of drones patrolling a 3D environment. Considering that the swarm configuration is significantly constrained by the geometry of the scene, experimental results show that it is possible to overcome the limitations of an ineffective configuration by exploiting camera-in-view side information.
Simone Milani, Alvise Memo
ICIP1
2016 A rate control algorithm for video coding in augmented reality applications
abstract
Most of latest-generation multimedia systems are equipped with increasingly-effective object detection algorithms (e.g., intelligent video surveillance systems, augmented reality applications, sharing platforms for multimedia data, etc.). Unfortunately, image and video compression makes object detection more difficult since such operations erase most of the computed visual features. In this paper we propose a video rate control strategy that exploits a saliency metric to identify image regions where features are likely to be found. According to the pixel statistics, the approach decides whether to increase the coding quality or include some side information that specifies the keypoint locations. Experimental results on HEVC coder show that the proposed rate control algorithm improves both the detection accuracy and the rate-distortion performance with respect to state-of-the-art strategies.
Simone Milani, Gianluca Agresti, Giancarlo Calvagno
PCS1
2016 Compression of multiple user photo galleries
Simone Milani
Image Vis. Comput.1
2016 Correction and interpolation of depth maps from structured light infrared sensors
Simone Milani, Giancarlo Calvagno
Signal Process. Image Commun.1
2016 Codec and GOP Identification in Double Compressed Videos
abstract
Video content is routinely acquired and distributed in a digital compressed format. In many cases, the same video content is encoded multiple times. This is the typical scenario that arises when a video, originally encoded directly by the acquisition device, is then re-encoded, either after an editing operation, or when uploaded to a sharing website. The analysis of the bitstream reveals details of the last compression step (i.e., the codec adopted and the corresponding encoding parameters), while masking the previous compression history. Therefore, in this paper, we consider a processing chain of two coding steps, and we propose a method that exploits coding-based footprints to identify both the codec and the size of the group of pictures (GOPs) used in the first coding step. This sort of analysis is useful in video forensics, when the analyst is interested in determining the characteristics of the originating source device, and in video quality assessment, since quality is determined by the whole compression history. The proposed method relies on the fact that lossy coding is an (almost) idempotent operation. That is, re-encoding a video sequence with the same codec and coding parameters produces a sequence that is similar to the former. As a consequence, if the second codec in the chain does not significantly alter the sequence, it is possible to analyze this sort of similarity to identify the first codec and the adopted GOP size. The method was extensively validated on a very large data set of video sequences generated by encoding content with a diversity of codecs (MPEG-2, MPEG-4, H.264/AVC, and DIRAC) and different encoding parameters. In addition, a proof of concept showing that the proposed method can also be used on videos downloaded from YouTube is reported.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
IEEE Trans. Image Process.2
2015 Three-dimensional reconstruction from heterogeneous video devices with camera-in-view information
abstract
The reconstruction of 3D scenes from unsynchronized and uncalibrated cameras has been a flourishing research area during the last years. In this work, a 3D modelization of the surrounding environment is enabled with an improvised ad-hoc camera networks of both static and mobile devices. The estimation can be significantly improved whenever one or more cameras (named here camera-in-views) can be localized within the field of view of other devices. Experimental results show that this additional information can improved the accuracy of the system up to 17 %.
Simone Milani
ICIP1
2015 Compression of photo collections using geometrical information
abstract
This paper proposes a novel scheme for the joint compression of photo collections framing the same object or scene. The proposed approach starts by locating corresponding features in the various images and then exploits a Structure from Motion algorithm to estimate the geometric relationships between the various images and their viewpoints. Then it uses 3D information and warping to predict images one from the other. Furthermore, graph algorithms are used to compute minimum weight topologies and identify the ordering of the input images that maximizes the efficiency of prediction. The obtained data is fed to a modified HEVC coder to perform the compression. Experimental results show that the proposed scheme outperforms competing solutions and can be efficiently employed for the storage of large image collections in the virtual exploration of architectural landmarks or in photo sharing websites.
Simone Milani, Pietro Zanuttigh
ICME1
2015 Performance evaluation of the 1st and 2nd generation Kinect for multimedia applications
abstract
Microsoft Kinect had a key role in the development of consumer depth sensors being the device that brought depth acquisition to the mass market. Despite the success of this sensor, with the introduction of the second generation, Microsoft has completely changed the technology behind the sensor from structured light to Time-Of-Flight. This paper presents a comparison of the data provided by the first and second generation Kinect in order to explain the achievements that have been obtained with the switch of technology. After an accurate analysis of the accuracy of the two sensors under different conditions, two sample applications, i.e., 3D reconstruction and people tracking, are presented and used to compare the performance of the two sensors.
Simone Zennaro, Matteo Munaro, Simone Milani, Pietro Zanuttigh, Andrea Bernardi, Stefano Ghidoni, Emanuele Menegatti
ICME3
2015 Distributed Multiple Description Video Transmission via Noncooperative Games With Opportunistic Players
abstract
Recent works have shown how multiple description coding proves to be an effective solution for multimedia streaming over peer-to-peer (P2P) and content delivery networks (CDNs). However, the presence of losses and congestions throughout the network affects the visual quality of the reconstructed sequence at the end terminal. These inconveniences can be mitigated by specifying different levels of Quality of Service, but an optimal packet classification is hard to obtain since P2P and CDN protocols operate at higher protocol layers (ignoring network conditions of the lowest stages), the network can be quite distributed and little information can be available regarding other network segments involved in the transmission. The peculiarities of the transmission scenario require a distributed and robust packet classification strategy that grants both intra-stream and inter-stream diversities among the loss patterns for the different streams. The classification approach presented here is characterized via a noncooperative game, where the different uploading nodes are players/descriptions competing for the allocation of the available network resources. Each player may switch from a selfish strategy to a more cooperative strategy according to its convenience (opportunistic players). Experimental results show that the proposed solution proves to be quite effective under different network scenarios.
Simone Milani, Giancarlo Calvagno
IEEE Trans. Circuits Syst. Video Technol.1
2014 Demosaicing strategy identification via eigenalgorithms
abstract
The identification of the camera that has acquired a specific image can be performed via several device-related footprints. Among these, it is possible to look for the traces left by the adopted color demosaicing strategy, which varies according to the camera model and vendor. The paper presents an identification strategy that re-processes the analyzed image with a set of distinctive CFA interpolation algorithms (eigenalgorithms) and, according to the correlation of the output with the original image, builds a set of features that permits identifying the algorithm. The proposed solution performs well with respect to other state-of-the-art solutions also when the analyzed image is severely compressed.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2014 Antiforensic synthesis of motion vectors using template algorithms
abstract
The identification of the video camera employed to acquire a video sequence is made possible by a large set of different footprints. Since video signals are always available in a compressed format, some of the most significant traces can be related to the coding tools of the implemented video codec (e.g., rate-distortion optimization, motion estimation strategy, etc.). As a matter of fact, an effective antiforensic attack, which aims at fooling the tools that identify the acquisition device, must appropriately alter these footprints. In the paper, we present an antiforensic strategy that targets a video camera detector which is based on the identification of the motion estimation strategy used by the video coder. The proposed approach synthesizes a set of motion vectors that approximate those that would have been generated by the algorithm to be mimicked. This method proves to be effective in attacking the detector while preserving the coding efficiency.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2014 Audio tampering detection using multimodal features
abstract
The authenticity verification of a User Generated Audio-Video content relative to a real event can be a very critical task especially when the content is shared on the Internet. Audio-Video files need to be checked in order to verify the origin of the content and the absence of alterations that could have changed their semantic content. The paper presents a multimodal approach for audio tampering detection that analyzes both the audio component and the video component of a recorded video file. The proposed solution estimates the volumetric characteristics of the environment where the multimedia content has been captured both from the video and audio signals. Then, the approach checks the consistency of the environment characteristics estimated from the audio signal with respect to those estimated from video files. The proposed solution proves to be useful in identifying video fakes and bootlegs, although it proves to be useful for the localization of added audio effects in a movie or radio track.
Simone Milani, Pier Francesco Piazza, Paolo Bestagini, Stefano Tubaro
ICASSP1
2014 Who is my parent? Reconstructing video sequences from partially matching shots
abstract
Nowadays, a significant fraction of the available video content is created by reusing already existing online videos. In these cases, the source video is seldom reused as is. Conversely, it is typically time clipped to extract only a subset of the original frames, and other transformations are commonly applied (e.g., cropping, logo insertion, etc.). In this paper, we analyze a pool of videos related to the same event or topic. We propose a method that aims at automatically reconstructing the content of the original source videos, i.e., the parent sequences, by splicing together sets of near-duplicate shots seemingly extracted from the same parent sequence. The result of the analysis shows how content is reused, thus revealing the intent of content creators, and enables us to reconstruct a parent sequence also when it is no longer available online. In doing so, we make use of a robust-hash algorithm that allows us to detect whether groups of frames are near-duplicates. Based on that, we developed an algorithm to automatically find near-duplicate matchings between multiple parts of multiple sequences. All the near-duplicate parts are finally temporally aligned to reconstruct the parent sequence. The proposed method is validated with both synthetic and real world datasets downloaded from YouTube.
Silvia Lameri, Paolo Bestagini, Andrea Melloni, Simone Milani, Anderson Rocha 0001, Marco Tagliasacchi, Stefano Tubaro
ICIP4
2013 Detection of temporal interpolation in video sequences
abstract
Nowadays, considering the availability of relatively cheap devices and powerful editing software, video tampering is a relatively easy task. Video sequences can be tampered with by performing, e.g., temporal splicing. However, if the sequences spliced together do not share the same frame rate, they have to be temporally interpolated beforehand. This operation is often made using motion compensated interpolators, which allow to minimize visual artifacts. In this paper we propose a detector of this kind of interpolation. Moreover, the detector is capable of identifying the interpolation factor used, allowing an analyst to uncover the original frame rate of a sequence. This method relies on the analysis of the correlation introduced by the filter adopted by the interpolator. Results show that detection is successful, provided that the number of observed interpolated frames is large enough. Moreover, tests on compressed sequences obtained from television broadcasts validate the method in a real world scenario.
Paolo Bestagini, S. Battaglia, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2013 A saliency-based rate control for people detection in video
abstract
Most of latest-generation multimedia systems are equipped with increasingly-effective object detection algorithms (e.g., intelligent video surveillance systems, augmented reality applications, sharing platforms for multimedia data, etc.). Unfortunately, images and video are usually available in compressed formats, which makes object detection more difficult because of the additional distortion noise. In this paper we show that it is possible to mitigate this problem by introducing a rate allocation algorithm that preserves important details for object identification algorithms. We propose a saliency map that identifies crucial elements for detectors. Then, we map saliency values to the value of the quantization parameter to be used by the video coder. Experimental results on HEVC coder show that the proposed rate control algorithm improves the accuracy with respect to the standard strategy.
Simone Milani, Riccardo Bernardini, Roberto Rinaldo
ICASSP1
2013 Antiforensics attacks to Benford's law for the detection of double compressed images
abstract
Researchers have been recently challenging the robustness of forensic algorithms by designing antiforensic strategies that try to fool them. In this paper, we propose an antiforensic strategy that targets double image compression detectors based on Benford's law (or first digit law). The proposed approach is able to modify the first digit statistics of the considered data (a double compressed image) to fool single/double compression detectors based on Benford's law. In this way, the proposed strategy tries to mimick the effects of a single compression with limited additional distortion. The presented algorithm performs better than previous state-of-the-art antiforensic strategies and can be easily extended to other fraud detection methods.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2013 No-reference quality metric for depth maps
abstract
The performances of several 3D imaging/video applications (going from 3DTV to video surveillance) benefit from the estimation or acquisition of accurate and high quality depth maps. However, the characteristics of depth information is strongly affected by the procedure employed in its acquisition or estimation (e.g., stereo evaluation, ToF cameras, structured light sensors, etc.), and the very definition of “quality” for a depth map is still under investigation. In this paper we proposed an unsupervised quality metric for depth information in Depth Image Based Rendering signals that predicts the accuracy in synthesizing 3D models and lateral views by using the considered depth information. The metric has been tested on depth maps generate with different algorithms and sensors. Moreover, experimental results show how it is possible to progressively improve the performance of 3D modelization by controlling the device/algorithm with this metric.
Simone Milani, Daniele Ferrario, Stefano Tubaro
ICIP1
2013 Identification of the motion estimation strategy using eigenalgorithms
abstract
The identification of the device, or device model, that was used to acquire a video sequence is a very challenging task, since it has to rely on subtle traces left by the processing steps applied to the raw acquired data. Previous works have tried to address this problem leveraging the traces left by the imaging sensor. However, in the case of video, lossy coding is often quite aggressive, thus making these methods impractical. In this work, we reverse the analysis strategy and exploit the traces left by lossy coding as telltale for the adopted acquisition device. Specifically, we aim at detecting the implementation of the video codec by identifying the adopted motion estimation algorithm. Indeed, motion estimation is not defined in video coding standards and, as such, it represents one of the non-normative tools that can be customized in the design of the encoder. The key tenet consists in studying the correlation between the motion vectors obtained from the decoded bitstream, and those computed using a set of known and diverse motion estimation algorithms, called eigenalgorithms. In our work, we generalize a method recently appeared in the literature, which assumes that the motion estimation algorithm used is necessarily one of those available during the analysis. Experimental results show that the approach is able to successfully identify the motion estimation algorithm in most cases.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICIP1
2013 Game-theoretic rate-distortion-complexity optimization for HEVC
abstract
This paper presents an algorithm for rate-distortion-complexity optimization for the emerging High Efficiency Video Coding (HEVC) standard, whose high computational requirements urge the need for low-complexity optimization algorithms. Optimization approaches need to specify different complexity profiles in order to tailor the computational load to the different hardware and power-supply resources of devices. In this work, we focus on optimizing the quantization parameter and partition depth in HEVC via a game-theoretic approach. The proposed rate control strategy alone provides 0.2 dB improvement compared to the approach implemented in HEVC reference software, while rate-distortion-complexity optimization allows very accurate complexity control providing at the same time rate-distortion performance close to the optimal one.
Anna Ukhanova, Simone Milani, Søren Forchhammer
ICIP2
2013 Local tampering detection in video sequences
abstract
Video sequences are often believed to provide stronger forensic evidence than still images, e.g., when used in lawsuits. However, a wide set of powerful and easy-to-use video authoring tools is today available to anyone. Therefore, it is possible for an attacker to maliciously forge a video sequence, e.g., by removing or inserting an object in a scene. These forms of manipulation can be performed with different techniques. For example, a portion of the original video may be replaced by either a still image repeated in time or, in more complex cases, by a video sequence. Moreover, the attacker might use as source data either a spatio-temporal region of the same video, or a region taken from an external sequence. In this paper we present the analysis of the footprints left when tampering with a video sequence, and propose a detection algorithm that allows a forensic analyst to reveal video forgeries and localize them in the spatio-temporal domain. With respect to the state-of-the-art, the proposed method is completely unsupervised and proves to be robust to compression. The algorithm is validated against a dataset of forged videos available online.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
MMSP2
2013 Distributed video coding based on lossy syndromes generated in hybrid pixel/transform domain
Simone Milani, Giancarlo Calvagno
Signal Process. Image Commun.1
2012 Video codec identification
abstract
Video content is routinely acquired and distributed in digital format. Therefore, it is customary to have the content encoded multiple times. In this paper we consider a processing chain of two coding steps and we propose a method that aims at identifying the type of codec used in the first step, by analyzing its coding-based footprints. The method relies on the fact that lossy coding is an almost idempotent operation, i.e., re-encoding the reconstructed sequence with the same codec and coding parameters produces a sequence that is highly correlated with the input one. As a consequence, it is possible to analyze this sort of correlation to identify the first codec provided that the second codec does not introduce severe quality degradation. The proposed solution finds several applications in the field of multi-media forensics, e.g. to identify the device that generated the original video stream or detect collages of different sequences.
Paolo Bestagini, Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP3
2012 Joint denoising and interpolation of depth maps for MS Kinect sensors
abstract
Infrared structured light sensors are widely employed for control applications, gaming, acquisition of dynamic and static 3D scenes. Recent developments have lead to the availability on the market of low-cost sensors which prove to be extremely sensitive to noise, light conditions, materials, the surface nature of the objects, and their distance from the camera. As a matter of fact, accurate denoising and interpolation strategies are needed. The paper presents a quality enhancement strategy for depth maps targeting low-cost IR structured light sensors. The approach has been tested using the MS Xbox Kinect device in both indoor and outdoor scenarios under different light conditions.
Simone Milani, Giancarlo Calvagno
ICASSP1
2012 Discriminating multiple JPEG compression using first digit features
abstract
The analysis of double-compressed images is a problem largely studied by the multimedia forensics community, as it might be exploited, e.g., for tampering localization or source device identification. In many practical scenarios, e.g. photos uploaded on blogs, on-line albums, and photo sharing Web sites, images might be compressed several times. However, the identification of the number of compression stages applied to an image remains an open issue. This paper proposes a forensic method based on the analysis of the distribution of the first significant digits of DCT coefficients, which is modeled according to Benford's law. The method relies on a set of Support Vector Machine (SVM) classifiers and allows us to accurately identify the number of compression stages applied to an image. Up to four consecutive compression stages were considered in the experimental validation. The proposed approach extends and outperforms the previously published methods aimed at detecting double JPEG compression.
Simone Milani, Marco Tagliasacchi, Stefano Tubaro
ICASSP1
2012 Adaptive denoising filtering for object detection applications
abstract
The widespread of augmented reality applications, cognitive video surveillance, autonomous or supportive navigation systems, has increased the importance of accurate object detection algorithms. However, the presence of noise depending on the characteristics of the acquisition device, on lighting intensity and directions, and on weather conditions, could severely degrade the performance of such applications. As a matter of fact, effective ad-hoc denoising strategies are required since traditional noise removal algorithms designed to improve the quality of the image, could even worsen the accuracy of detection. This paper presents a low-cost adaptive filtering strategy that adapts the characteristics of the filter depending on the impact of each image region on the feature sets. This solution permits improving the correct detection percentage of approximately 30%with respect to using noisy images. The approach is generally intended for object detection algorithms based on Histogram-of-Oriented-Gradients (HOG) and can run in real time on a limited complexity hardware.
Simone Milani, Riccardo Bernardini, Roberto Rinaldo
ICIP1
2012 Multiple compression detection for video sequences
abstract
Nowadays, thanks to the increasingly availability of powerful processors and user friendly applications, the editing of video sequences is becoming more and more frequent. Moreover, after each editing step, any video object is almost always encoded in order to store it using a less amount of memory. For this reason, inferring the number of compression steps that have been applied to such a multimedia object is an important clue in order to assess its authenticity. In this paper we propose a method to recover the number of compression steps applied to a video sequence. In order to accomplish this goal, we make use of a classifier based on multiple Support Vector Machines (SVM) exploiting the Benford's law. Indeed, the feature vectors used to train and test the SVM are based on the statistics of the most significant digit of quantized transform coefficients. The proposed method is tested with a generic hybrid video encoder combining motion-compensation and block coding. Results show that this method is able to discriminate up to three compression stages with high accuracy.
Simone Milani, Paolo Bestagini, Marco Tagliasacchi, Stefano Tubaro
MMSP1
2011 Segmentation-based motion compensation for enhanced video coding
abstract
Traditional video compression is based on block-based motion estimation and compensation. However, experimental results and the latest standards have proved that adapting the size and the shape of the motion estimation area to the objects in the scene significantly improves the performance of the overall coding process. The current paper presents a new segmentation-based coding strategy that partitions the input frame into irregularly-shaped regions. These segments are then used to estimate the motion-compensated prediction in the previous frames. The scheme outperforms the rate-distortion performance of H.264/AVC of 2 dB with a reasonable increment of the encoding complexity due to segmentation.
Simone Milani, Giancarlo Calvagno
ICIP1
2011 Efficient depth map compression exploiting segmented color data
abstract
3D video representations usually associate to each view a depth map with the corresponding geometric information. Many compression schemes have been proposed for multi-view video and for depth data, but the exploitation of the correlation between the two representations to enhance compression performances is still an open research issue. This paper presents a novel compression scheme that exploits a segmentation of the color data to predict the shape of the different surfaces in the depth map. Then each segment is approximated with a parameterized plane. In case the approximation is sufficiently accurate for the target bit rate, the surface coefficients are compressed and transmitted. Otherwise, the region is coded using a standard H.264/AVC Intra coder. Experimental results show that the proposed scheme permits to outperformthe standardH.264/AVC Intra codec on depth data and can be effectively included into multi-view plus depth compression schemes.
Simone Milani, Pietro Zanuttigh, Marco Zamarin, Søren Forchhammer
ICME1
2011 3DTV streaming over Peer-to-Peer networks using FEC-based noncooperative multiple description
abstract
Recent works have shown how Multiple Description Coding (MDC) and Peer-to-Peer transmission protocols prove to be extremely helpful in streaming video sequences. The quality of the reconstructed sequence can be significantly improved by differentiating the Quality-of-Service levels via distributed packet classifications performed by each peer independently. The packet labelling can be modelled via noncooperative games where the different uploading nodes are players/descriptions competing for the available network resources. The paper shows how this framework can be effectively employed for the transmission of 3D video sequences adopting an FEC-based multiple description scheme that encapsulates different kinds of data streams (color information, geometry) into a set of equally-important packets. Experimental results show that the classification strategy based on game theory improves the quality of the reconstructed sequence with respect to standard classification strategies.
Simone Milani, Marco Gaggio, Giancarlo Calvagno
ISCC1
2011 Resolution Scalable Image Coding With Reversible Cellular Automata
abstract
In a resolution scalable image coding algorithm, a multiresolution representation of the data is often obtained using a linear filter bank. Reversible cellular automata have been recently proposed as simpler, nonlinear filter banks that produce a similar representation. The original image is decomposed into four subbands, such that one of them retains most of the features of the original image at a reduced scale. In this paper, we discuss the utilization of reversible cellular automata and arithmetic coding for scalable compression of binary and grayscale images. In the binary case, the proposed algorithm that uses simple local rules compares well with the JBIG compression standard, in particular for images where the foreground is made of a simple connected region. For complex images, more efficient local rules based upon the lifting principle have been designed. They provide compression performances very close to or even better than JBIG, depending upon the image characteristics. In the grayscale case, and in particular for smooth images such as depth maps, the proposed algorithm outperforms both the JBIG and the JPEG2000 standards under most coding conditions.
Lorenzo Cappellari, Simone Milani, Carlos Cruz-Reyes, Giancarlo Calvagno
IEEE Trans. Image Process.2
2011 Fast H.264/AVC FRExt Intra Coding Using Belief Propagation
abstract
In the H.264/AVC FRExt coder, the coding performance of Intra coding significantly overcomes the previous still image coding standards, like JPEG2000, thanks to a massive use of spatial prediction. Unfortunately, the adoption of an extensive set of predictors induces a significant increase of the computational complexity required by the rate-distortion optimization routine. The paper presents a complexity reduction strategy that aims at reducing the computational load of the Intra coding with a small loss in the compression performance. The proposed algorithm relies on selecting a reduced set of prediction modes according to their probabilities, which are estimated adopting a belief-propagation procedure. Experimental results show that the proposed method permits saving up to 60 % of the coding time required by an exhaustive rate-distortion optimization method with a negligible loss in performance. Moreover, it permits an accurate control of the computational complexity unlike other methods where the computational complexity depends upon the coded sequence.
Simone Milani
IEEE Trans. Image Process.1
2011 A cognitive approach for effective coding and transmission of 3D video
abstract
Future multimedia applications will rely on the transmission of 3D video contents within heterogeneous fruition scenarios, and as a matter of fact, the reliable delivery of 3D video signals proves to be a crucial issue in such communications. To this purpose, multimedia communication experts have been designing cross-layer strategies to improve the quality of the perceived 3D experience. This article presents a new cross-layer strategy, called Cognitive Source Coding (CSC), that defines a new 3D video system able to identify the different elements of the 3D scene and choose the most appropriate coding strategy.
Simone Milani, Giancarlo Calvagno
ACM Trans. Multim. Comput. Commun. Appl.1
2010 A cognitive approach for effective coding and transmission of 3D video
abstract
Reliable delivery of 3D video contents to a wide set of users is expected to be the next big revolution in multimedia applications provided that it is possible to grant a certain level of Quality-of-Experience (QoE) to the end user.
Simone Milani, Giancarlo Calvagno
ACM Multimedia1
2010 A novel multi-view image coding scheme based on view-warping and 3D-DCT
Marco Zamarin, Simone Milani, Pietro Zanuttigh, Guido M. Cortelazzo
J. Vis. Commun. Image Represent.2
2010 Multiple Description Distributed Video Coding Using Redundant Slices and Lossy Syndromes
abstract
During the last years, video coding designers have proposed robust coding approaches that combine Multiple Description Coding (MDC) schemes with Distributed Video Coding (DVC) principles. In this way, it is possible to obtain a better error resilience since the distortion drifting through the sequence is significantly mitigated by the DVC coding unit. The paper presents a Multiple Description Distributed Video Coder (MDDVC) that codes the input video signal generating a set of "lossy" syndromes for each pixel block and creates different descriptions multiplexing primary and redundant video packets. Experimental results show that at high loss probabilities the proposed solution improves the results of the original MDC approach.
Simone Milani, Giancarlo Calvagno
IEEE Signal Process. Lett.1
2010 A Depth Image Coder Based on Progressive Silhouettes
abstract
An efficient compression of depth maps proves to be a crucial element in the transmission and storage of 3-D scenes. However, the peculiarities of geometry information make the traditional coding paradigms for natural images less effective for the coding of depth images. The letter presents a novel coding scheme that employs an oversegmentation of the input depth image into a huge set of small regions. These regions are then fused together according to the target number of objects that the algorithm needs to identify in the representation. This procedure is iterated more than once generating several refinement layers that permit obtaining a progressively-increasing quality in the scene. Experimental results show that in most cases the proposed approach reaches a better coding performance with respect to previous coding methods.
Simone Milani, Giancarlo Calvagno
IEEE Signal Process. Lett.1
2009 A Binary Image Scalable Coder Based on Reversible Cellular Automata Transform and Arithmetic Coding
abstract
In this work, we have designed an efficient arithmetic coder for the non-linear bi-level image coder based on reversible cellular automata transform reported in the work of Cruz-Reyes and Kari (2008). The proposed approach relies on non-linear transform operations based on RmbCA, decomposing the image into four subimages (named LL, LH, HL, HH in lexicographic order) in such a way that the LL subimage represents a low resolution version of the original image. Performing an edge detection on this subimage, it is possible to identify the coordinates where there is a high probability that the other subimages present a black pixels. In this way it is possible to minimize the amount of data that have to be processed reducing the coded bit stream. In order to increase the probability that black pixels lie in the selected region, the coding algorithm enlarges the identified region according to the values of Sobel operators computed on the LL subimage.
Simone Milani, Carlos Cruz-Reyes, Jarkko Kari 0001, Giancarlo Calvagno
DCC1
2009 A game theory based classification for distributed downloading of multiple description coded video
abstract
Recent works have shown how Multiple Description Coding (MDC) proves to be an effective solution for the multimedia transmission over Peer-to-Peer and Content Delivery Networks. However, despite the flexibility and scalability of these novel transmission paradigms, packet streams are still affected by losses, and an effective packet labelling that differentiates the QoS levels for the transmitted data is required. The paper presents a Game Theory based approach to classify the MDC coded video packets that are downloaded from a set of terminals spread throughout the network. The proposed approach improves the quality of the reconstructed sequence with limited feedback information from the network.
Simone Milani, Giancarlo Calvagno
ICIP1
2009 A low-complexity rate allocation algorithm for joint source-channel video coding
Simone Milani, Giancarlo Calvagno
Signal Process. Image Commun.1
2009 A Low-Complexity Cross-Layer Optimization Algorithm for Video Communication Over Wireless Networks
abstract
Recent years have witnessed a rapid increment in video applications over wireless networks including on-demand video streaming and video phoning. This growth has also brought the need to find a good compromise in the conflict between resource limitations affecting mobile devices and the desire for high-quality multimedia services. It is possible to face this problem adopting a cross-layer strategy that jointly tunes the parameters of each layer in the network protocol stack. In this optimization strategy, complexity is one of the most significant issues because of the limited computational resources and power supply. The paper presents a low-complexity cross-layer algorithm that is able to jointly tune the parameters of different protocol layers by adopting simple but effective models. The quality of the reconstructed video sequence, the produced bit rate, and the service class associated to each packet are seen as functions of the percentage of null DCT coefficients. This modeling permits to find a closed-form solution to the joint optimization problem that can be computed with a limited number of operations and grants, at the same time, a good visual quality in the reconstructed sequence.
Simone Milani, Giancarlo Calvagno
IEEE Trans. Multim.1
2008 A low-complexity packet classification algorithm for multiple description video streaming over IEEE802.11E networks
abstract
The robust transmission of video sequences over wireless LANs presents several challenging problems concerning the presence of packet losses, delays, and bandwidth limitations. Their effect on the visual quality of the sequence reconstructed by the end user can be mitigated by adopting a wireless architecture that is able to support different levels of quality-of-service (QoS), like the IEEE 802.11e standard, and by compressing the video sequence to be transmitted using a robust source coder, like a multiple description coding (MDC) scheme. This paper presents a MDC-based video streaming architecture that tries to adaptively optimize the performance of both solutions by assigning the RTP video packets produced by the MDC encoder to the different QoS classes of IEEE 802.11e. Simulation results show that the performance is significantly improved with respect to a non-adaptive solution.
Simone Milani, Giancarlo Calvagno, Riccardo Bernardini, Roberto Rinaldo
ICIP1
2008 An Accurate Low-Complexity Rate Control Algorithm Based on (rho, Eq)-Domain
abstract
The standard H.264/AVC defines an efficient coding architecture both for coding applications where bandwidth or storage capacity is limited (e.g., video telephony or video conferencing over mobile channels and devices) and for applications that require high reconstruction quality and bit rate (e.g., HDTV). Since its main applications concern video communication over time-varying bandwidth channels, the bit rate has to be controlled with scalable algorithms that can be implemented on low resource devices. The paper describes a rate control algorithm that needs reduced memory area and complexity compared to other ones. The number of coded bits for each frame is accurately predicted through the percentage of null quantized transform coefficients, which is related to the quantization step via the energy of the quantized signal. It is possible to design a rate control algorithm based on this model that provides a good compression performance at a low computational cost.
Simone Milani, Luca Celetto, Gian Antonio Mian
IEEE Trans. Circuits Syst. Video Technol.1
2007 Achieving H.264-like compression efficiency with distributed video coding
abstract
Recently, a new class of distributed source coding (DSC) based video coders has been proposed to enable low-complexity encoding. However, to date, these low-complexity DSC-based video encoders have been unable to compress as efficiently as motion-compensated predictive coding based video codecs, such as H.264/AVC, due to insufficiently accurate modeling of video data. In this work, we examine achieving H.264-like high compression efficiency with a DSC-based approach without the encoding complexity constraint. The success of H.264/AVC highlights the importance of accurately modeling the highly non-stationary video data through fine-granularity motion estimation. This motivates us to deviate from the popular approach of approaching the Wyner-Ziv bound with sophisticated capacity-achieving channel codes that require long block lengths and high decoding complexity, and instead focus on accurately modeling video data. Such a DSC-based, compression-centric encoder is an important first step towards building a robust DSC-based video coding framework.
Simone Milani, Kannan Ramchandran
VCIP1