VLDB 2026 Research / reviewers in the wild / expert
Pier Luigi Dragotti
dblp:47/1135
· DBLP profile ↗
141ranked-venue papers
17as first author
27since 2021 · last 2025
0000-0002-6073-2807ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 120 · 14 first-author · 17 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Theory of computation · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LotteryCodec: Searching the Implicit Representation in a Random Network for Low-Complexity Image CompressionabstractWe introduce and validate the lottery codec hypothesis, which states that untrained subnetworks within randomly initialized networks can serve as synthesis networks for overfitted image compression, achieving rate-distortion (RD) performance comparable to trained networks. This hypothesis leads to a new paradigm for image compression by encoding image statistics into the network substructure. Building on this hypothesis, we propose LotteryCodec, which overfits a binary mask to an individual image, leveraging an over-parameterized and randomly initialized network shared by the encoder and the decoder. To address over-parameterization challenges and streamline subnetwork search, we develop a rewind modulation mechanism that improves the RD performance. LotteryCodec outperforms VTM and sets a new state-of-the-art in single-image compression. LotteryCodec also enables adaptive decoding complexity through adjustable mask ratios, offering flexible compression solutions for diverse device constraints and application requirements. Gongpu Chen, Pier Luigi Dragotti, Deniz Gündüz |
ICML | 3 |
| 2025 | Tracing the Roots: Leveraging Temporal Dynamics in Diffusion Trajectories for Origin AttributionabstractDiffusion models have transformed image synthesis through iterative denoising, by defining trajectories from noise to coherent data. While their capabilities are widely celebrated, a critical challenge remains unaddressed: ensuring responsible use by verifying whether an image originates from a model's training set, its novel generations or external sources. We introduce a framework that analyzes diffusion trajectories for this purpose. Specifically, we demonstrate that temporal dynamics across the entire trajectory allow for more robust classification and challenge the widely-adopted "Goldilocks zone" conjecture, which posits that membership inference is effective only within narrow denoising stages. More fundamentally, we expose critical flaws in current membership inference practices by showing that representative methods fail under distribution shifts or when model-generated data is present. For model attribution, we demonstrate a first white-box approach directly applicable to diffusion. Ultimately, we propose the unification of data provenance into a single, cohesive framework tailored to modern generative systems. Andreas Floros 0002, Seyed-Mohsen Moosavi-Dezfooli, Pier Luigi Dragotti |
NeurIPS | 3 |
| 2025 | Enhanced accuracy in first-spike coding using current-based adaptive LIF neuronabstractFirst spike timings are crucial for decision-making in spiking neural networks (SNNs). A recently introduced first-spike (FS) coding method demonstrates comparable accuracy to firing-rate (FR) coding in processing complex temporal information through supervised learning. However, its performance still falls behind advanced approaches. In order to explore the potential of FS coding, we enhance the capability of SNNs in classifying auditory datasets by improving neural dynamics. We propose a current-based adaptive LIF neuron (CuAdLIF) with delayed responses and membrane potential adaptation to enhance temporal correlations and preserve long short-memory. Furthermore, we introduce strategies to minimize delays in decision-making and enable adaptive training for FS coding. Results show that the CuAdLIF neuron enhances the extraction of temporal features and significantly improves FS coding accuracy. In addition, our strategies effectively reduce output time delays. Pier Luigi Dragotti |
Neural Networks | 2 |
| 2025 | A Lightweight Deep Exclusion Unfolding Network for Single Image Reflection RemovalabstractSingle Image Reflection Removal (SIRR) is a canonical blind source separation problem and refers to the issue of separating a reflection-contaminated image into a transmission and a reflection image. The core challenge lies in minimizing the commonalities among different sources. Existing deep learning approaches either neglect the significance of feature interactions or rely on heuristically designed architectures. In this paper, we propose a novel Deep Exclusion unfolding Network (DExNet), a lightweight, interpretable, and effective network architecture for SIRR. DExNet is principally constructed by unfolding and parameterizing a simple iterative Sparse and Auxiliary Feature Update (i-SAFU) algorithm, which is specifically designed to solve a new model-based SIRR optimization formulation incorporating a general exclusion prior. This general exclusion prior enables the unfolded SAFU module to inherently identify and penalize commonalities between the transmission and reflection features, ensuring more accurate separation. The principled design of DExNet not only enhances its interpretability but also significantly improves its performance. Comprehensive experiments on four benchmark datasets demonstrate that DExNet achieves state-of-the-art visual and quantitative results while utilizing only approximately 8% of the parameters required by leading methods. Junjie Huang 0001, Tianrui Liu 0001, Xinwang Liu 0002, Meng Wang 0001, Pier Luigi Dragotti |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | CommIN: Semantic Image Communications as an Inverse Problem with INN-Guided Diffusion ModelsabstractJoint source-channel coding schemes based on deep neural networks (DeepJSCC) have recently achieved remarkable performance for wireless image transmission. However, these methods usually focus only on the distortion of the reconstructed signal at the receiver side with respect to the source at the transmitter side, rather than the perceptual quality of the reconstruction which carries more semantic information. As a result, severe perceptual distortion can be introduced under extreme conditions such as low bandwidth and low signal-to-noise ratio. In this work, we propose CommIN, which views the recovery of high-quality source images from degraded reconstructions as an inverse problem. To address this, CommIN combines Invertible Neural Networks (INN) with diffusion models, aiming for superior perceptual quality. Through experiments, we show that our CommIN significantly improves the perceptual quality compared to DeepJSCC under extreme conditions and outperforms other inverse problem approaches used in DeepJSCC. Jiakang Chen, Di You, Deniz Gündüz, Pier Luigi Dragotti |
ICASSP | 4 |
| 2024 | DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint PriorabstractSingle image reflection removal (SIRR) problem can be interpreted as a canonical blind source separation problem and is highly ill-posed. A parameter effective, fast learning and interpretable reflection removal algorithm is essential for many vision analysis applications. In this paper, we propose a novel model-inspired and learning-based SIRR method called Deep Unfolded Reflection Removal Network (DURRNet). It combines the merits of both model-based and learning-based paradigms, leading to a more interpretable and effective deep architecture. To achieve this, we first propose a model-based optimization approach and then obtain DURRNet by unfolding an iterative step into a Unfolded Separation Block (USB) based on proximal gradient descent. Key features of DURR-Net include the use of Invertible Neural Networks to impose the transform-based exclusion prior on the basis of natural image prior, as well as a coarse-to-fine architecture to fine-grain the reflection removal process. Extensive experiments on public datasets demonstrate that DURRNet achieves state-of-the-art results not only visually, quantitatively, but also effectively. Junjie Huang 0001, Tianrui Liu 0001, Jingyuan Xia, Meng Wang 0001, Pier Luigi Dragotti |
ICASSP | 5 |
| 2024 | Model-Based Explainable Deep Learning for Light-Field Microscopy ImagingabstractIn modern neuroscience, observing the dynamics of large populations of neurons is a critical step of understanding how networks of neurons process information. Light-field microscopy (LFM) has emerged as a type of scanless, high-speed, three-dimensional (3D) imaging tool, particularly attractive for this purpose. Imaging neuronal activity using LFM calls for the development of novel computational approaches that fully exploit domain knowledge embedded in physics and optics models, as well as enabling high interpretability and transparency. To this end, we propose a model-based explainable deep learning approach for LFM. Different from purely data-driven methods, the proposed approach integrates wave-optics theory, sparse representation and non-linear optimization with the artificial neural network. In particular, the architecture of the proposed neural network is designed following precise signal and optimization models. Moreover, the network's parameters are learned from a training dataset using a novel training strategy that integrates layer-wise training with tailored knowledge distillation. Such design allows the network to take advantage of domain knowledge and learned new features. It combines the benefit of both model-based and learning-based methods, thereby contributing to superior interpretability, transparency and performance. By evaluating on both structural and functional LFM data obtained from scattering mammalian brain tissues, we demonstrate the capabilities of the proposed approach to achieve fast, robust 3D localization of neuron sources and accurate neural activity identification. Pingfan Song, Herman Verinaz-Jadan, Carmel L. Howe, Amanda J. Foust, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 5 |
| 2023 | Sparse Asynchronous Samples from Networks of Tems for Reconstruction of Classes of Non-Bandlimited SignalsabstractWe present a signal driven multi-channel time encoding system for sampling signals with finite rate of innovation (FRI). The system produces samples in the form of finite differences from which the input signal can be exactly reconstructed. The use of finite differences allows diversity of TEM parameters between channels, particularly with regards to resetting characteristics and the multi-channel approach allows for better control and minimisation of the number of samples generated meaning the sampler can be considered more energy efficient than existing methods which inherently produce large numbers of samples. Furthermore, the system only generates samples in response to innovations in the signal. This makes the system well suited to sampling truly asynchronous signals consisting of bursts of activity interspersed with long indeterminate periods of quiescence while also generating sparse sets of samples. Marek Hilton, Pier Luigi Dragotti |
ICASSP | 2 |
| 2023 | Super-Resolution for Macro X-Ray Fluorescence Data Collected from Old Master PaintingsabstractMacro X-ray fluorescence (MA-XRF) scanning is commonly used to non-invasively analyse Old Master paintings by mapping the distribution of the chemical elements present in the artworks. The visual quality of the element distribution maps is very important for characterising the materials and understanding the execution and condition of the painting. However, this quality is limited by the acquisition time for the XRF datacube, resulting in a trade-off between signal-to-noise ratio (SNR) and spatial resolution. To solve this problem we propose to enhance the spatial resolution of the XRF datacube of a painting leveraging a corresponding high-resolution (HR) RGB image. We achieve that by introducing a method based on coupled dictionary learning along with a similarity constraint based on mutual information. In particular, we divide the RGB image and the XRF datacube into a common part and a unique part based on whether the information is shared or not, and then transfer the HR information between the two common parts, resulting in high-quality reconstructions. Numerical results show that our XRF super-resolution method outperforms the other state-of-the-art approaches. Su Yan 0003, Herman Verinaz-Jadan, Junjie Huang 0001, Nathan Daly, Catherine Higgitt, Pier Luigi Dragotti |
ICASSP | 6 |
| 2023 | INDigo: An INN-Guided Probabilistic Diffusion Algorithm for Inverse ProblemsabstractRecently it has been shown that using diffusion models for inverse problems can lead to remarkable results. However, these approaches require a closed-form expression of the degradation model and can not support complex degradations. To overcome this limitation, we propose a method (INDigo) that combines invertible neural networks (INN) and diffusion models for general inverse problems. Specifically, we train the forward process of INN to simulate an arbitrary degradation process and use the inverse as a reconstruction process. During the diffusion sampling process, we impose an additional data-consistency step that minimizes the distance between the intermediate result and the INN-optimized result at every iteration, where the INN-optimized image is composed of the coarse information given by the observed degraded image and the details generated by the diffusion process. With the help of INN, our algorithm effectively estimates the details lost in the degradation process and is no longer limited by the requirement of knowing the closed-form expression of the degradation model. Experiments demonstrate that our algorithm obtains competitive results compared with recently leading methods both quantitatively and visually. Moreover, our algorithm performs well on more complex degradation models and real-world low-quality images. Di You, Andreas Floros 0002, Pier Luigi Dragotti |
MMSP | 3 |
| 2023 | Generative Joint Source-Channel Coding for Semantic Image TransmissionabstractRecent works have shown that joint source-channel coding (JSCC) schemes using deep neural networks (DNNs), called DeepJSCC, provide promising results in wireless image transmission. However, these methods mostly focus on the distortion of the reconstructed signals with respect to the input image, rather than their perception by humans. However, focusing on traditional distortion metrics alone does not necessarily result in high perceptual quality, especially in extreme physical conditions, such as very low bandwidth compression ratio (BCR) and low signal-to-noise ratio (SNR) regimes. In this work, we propose two novel JSCC schemes that leverage the perceptual quality of deep generative models (DGMs) for wireless image transmission, namely InverseJSCC and GenerativeJSCC. While the former is an inverse problem approach to DeepJSCC, the latter is an end-to-end optimized JSCC scheme. In both, we optimize a weighted sum of mean squared error (MSE) and learned perceptual image patch similarity (LPIPS) losses, which capture more semantic similarities than other distortion metrics. InverseJSCC performs denoising on the distorted reconstructions of a DeepJSCC model by solving an inverse optimization problem using the pre-trained style-based generative adversarial network (StyleGAN). Our simulation results show that InverseJSCC significantly improves the state-of-the-art DeepJSCC in terms of perceptual quality in edge cases. In GenerativeJSCC, we carry out end-to-end training of an encoder and a StyleGAN-based decoder, and show that GenerativeJSCC significantly outperforms DeepJSCC both in terms of distortion and perceptual quality. Ece Naz Erdemir, Tze-Yang Tung, Pier Luigi Dragotti, Deniz Gündüz |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Sensing Diversity and Sparsity Models for Event Generation and Video Reconstruction from EventsabstractEvents-to-video (E2V) reconstruction and video-to-events (V2E) simulation are two fundamental research topics in event-based vision. Current deep neural networks for E2V reconstruction are usually complex and difficult to interpret. Moreover, existing event simulators are designed to generate realistic events, but research on how to improve the event generation process has been so far limited. In this paper, we propose a light, simple model-based deep network for E2V reconstruction, explore the diversity for adjacent pixels in V2E generation, and finally build a video-to-events-to-video (V2E2V) architecture to validate how alternative event generation strategies improve video reconstruction. For the E2V reconstruction, we model the relationship between events and intensity using sparse representation models. A convolutional ISTA network (CISTA) is then designed using the algorithm unfolding strategy. Long short-term temporal consistency (LSTC) constraints are further introduced to enhance the temporal coherence. In the V2E generation, we introduce the idea of having interleaved pixels with different contrast threshold and lowpass bandwidth and conjecture that this can help extract more useful information from intensity. Finally, V2E2V architecture is used to verify the effectiveness of this strategy. Results highlight that our CISTA-LSTC network outperforms state-of-the-art methods and achieves better temporal consistency. Sensing diversity in event generation reveals more fine details and this leads to a significantly improved reconstruction quality. Pier Luigi Dragotti |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Privacy-Aware Communication over a Wiretap Channel with Generative NetworksabstractWe study privacy-aware communication over a wiretap channel using end-to-end learning. Alice wants to transmit a source signal to Bob over a binary symmetric channel, while passive eavesdropper Eve tries to infer some sensitive attribute of Alice’s source based on its overheard signal. Since we usually do not have access to true distributions, we propose a data-driven approach using variational autoencoder (VAE)-based joint source channel coding (JSCC). We show through simulations with the colored MNIST dataset that our approach provides high reconstruction quality at the receiver while confusing the eavesdropper about the latent sensitive attribute, which consists of the color and thickness of the digits. Finally, we consider a parallel-channel scenario, and show that our approach arranges the information transmission such that the channels with higher noise levels at the eavesdropper carry the sensitive information, while the non-sensitive information is transmitted over more vulnerable channels. Ece Naz Erdemir, Pier Luigi Dragotti, Deniz Gündüz |
ICASSP | 2 |
| 2022 | Convolutional ISTA Network with Temporal Consistency Constraints for Video Reconstruction from Event CamerasabstractEvent cameras produce streams of events with high temporal resolution which do not suffer from motion blur. Current deep networks achieve high-quality video reconstruction from events, but most of them are large and difficult to interpret. In this work, we present a solution to this problem by systematically designing a deep network based on sparse representation. First, we investigate the relationship between events and intensity images. The reconstruction problem is then modelled as a sparse coding problem, which can be solved by the iterative shrinkage thresholding algorithm (ISTA). Second, we expand this into a convolutional ISTA network (CISTA) using algorithm unfolding. Finally, we introduce recurrent units and temporal similarity constraints to enhance the temporal consistency (TC) reconstruction of long videos. Results show that our CISTA-TC network achieves high-quality reconstruction compared with state-of-the-art methods, whilst leading to low memory consumption. Roxana Alexandru, Pier Luigi Dragotti |
ICASSP | 3 |
| 2022 | Perfect Reconstruction of Classes of Non-Bandlimited Signals from Projections with Unknown AnglesabstractIn this paper, we consider the 2D tomography problem for a finite number of point sources, where the line integral projections are taken at unknown angles. We address the problem of recovering the point sources and estimating the projection angles. Using the property of the Radon transform of a point source, which is a signal with Finite Rate of Innovation, we retrieve the projections using the annihilating filter method. The reconstruction method we propose is then able to unveil the 2D geometric information of the projection angles, as well as the locations of the point sources. Finally, we extend the approach to planar polygons. Renke Wang, Roxana Alexandru, Pier Luigi Dragotti |
ICASSP | 3 |
| 2022 | AI-Based Reconstruction for Fast MRI - A Systematic Review and Meta-AnalysisabstractCompressed sensing (CS) has been playing a key role in accelerating the magnetic resonance imaging (MRI) acquisition process. With the resurgence of artificial intelligence, deep neural networks and CS algorithms are being integrated to redefine the state of the art of fast MRI. The past several years have witnessed substantial growth in the complexity, diversity, and performance of deep-learning-based CS techniques that are dedicated to fast MRI. In this meta-analysis, we systematically review the deep-learning-based CS techniques for fast MRI, describe key model designs, highlight breakthroughs, and discuss promising directions. We have also introduced a comprehensive analysis framework and a classification system to assess the pivotal role of deep learning in CS-based acceleration for MRI. Carola-Bibiane Schönlieb, Pietro Liò, Tim Leiner, Pier Luigi Dragotti, Ge Wang 0001, Daniel Rueckert, David N. Firmin, Guang Yang 0006 |
Proc. IEEE | 5 |
| 2022 | Multi-Modal Convolutional Dictionary LearningabstractConvolutional dictionary learning has become increasingly popular in signal and image processing for its ability to overcome the limitations of traditional patch-based dictionary learning. Although most studies on convolutional dictionary learning mainly focus on the unimodal case, real-world image processing tasks usually involve images from multiple modalities, e.g., visible and near-infrared (NIR) images. Thus, it is necessary to explore convolutional dictionary learning across different modalities. In this paper, we propose a novel multi-modal convolutional dictionary learning algorithm, which efficiently correlates different image modalities and fully considers neighborhood information at the image level. In this model, each modality is represented by two convolutional dictionaries, in which one dictionary is for common feature representation and the other is for unique feature representation. The model is constrained by the requirement that the convolutional sparse representations (CSRs) for the common features should be the same across different modalities, considering that these images are captured from the same scene. We propose a new training method based on the alternating direction method of multipliers (ADMM) to alternatively learn the common and unique dictionaries in the discrete Fourier transform (DFT) domain. We show that our model converges in less than 20 iterations between the convolutional dictionary updating and the CSRs calculation. The effectiveness of the proposed dictionary learning algorithm is demonstrated on various multimodal image processing tasks, achieves better performance than both dictionary learning methods and deep learning based methods with limited training data. Fangyuan Gao, Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 5 |
| 2022 | WINNet: Wavelet-Inspired Invertible Network for Image DenoisingabstractImage denoising aims to restore a clean image from an observed noisy one. Model-based image denoising approaches can achieve good generalization ability over different noise levels and are with high interpretability. Learning-based approaches are able to achieve better results, but usually with weaker generalization ability and interpretability. In this paper, we propose a wavelet-inspired invertible network (WINNet) to combine the merits of the wavelet-based approaches and learning-based approaches. The proposed WINNet consists of K -scale of lifting inspired invertible neural networks (LINNs) and sparsity-driven denoising networks together with a noise estimation network. The network architecture of LINNs is inspired by the lifting scheme in wavelets. LINNs are used to learn a non-linear redundant transform with perfect reconstruction property to facilitate noise removal. The denoising network implements a sparse coding process for denoising. The noise estimation network estimates the noise level from the input image which will be used to adaptively adjust the soft-thresholds in LINNs. The forward transform of LINNs produces a redundant multi-scale representation for denoising. The denoised image is reconstructed using the inverse transform of LINNs with the denoised detail channels and the original coarse channel. The simulation results show that the proposed WINNet method is highly interpretable and has strong generalization ability to unseen noise levels. It also achieves competitive results in the non-blind/blind image denoising and in image deblurring. Junjie Huang 0001, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2022 | Mixed X-Ray Image Separation for Artworks With Concealed DesignsabstractIn this paper, we focus on X-ray images (X-radiographs) of paintings with concealed sub-surface designs (e.g., deriving from reuse of the painting support or revision of a composition by the artist), which therefore include contributions from both the surface painting and the concealed features. In particular, we propose a self-supervised deep learning-based image separation approach that can be applied to the X-ray images from such paintings to separate them into two hypothetical X-ray images. One of these reconstructed images is related to the X-ray image of the concealed painting, while the second one contains only information related to the X-ray image of the visible painting. The proposed separation network consists of two components: the analysis and the synthesis sub-networks. The analysis sub-network is based on learned coupled iterative shrinkage thresholding algorithms (LCISTA) designed using algorithm unrolling techniques, and the synthesis sub-network consists of several linear mappings. The learning algorithm operates in a totally self-supervised fashion without requiring a sample set that contains both the mixed X-ray images and the separated ones. The proposed method is demonstrated on a real painting with concealed content, Do na Isabel de Porcel by Francisco de Goya, to show its effectiveness. Junjie Huang 0001, Barak Sober, Nathan Daly, Catherine Higgitt, Ingrid Daubechies, Pier Luigi Dragotti, Miguel R. D. Rodrigues |
IEEE Trans. Image Process. | 7 |
| 2021 | Active Privacy-Utility Trade-Off Against A Hypothesis Testing AdversaryabstractWe consider a user releasing her data containing some personal information in return of a service. We model user’s personal information as two correlated random variables, one of them, called the secret variable, is to be kept private, while the other, called the useful variable, is to be disclosed for utility. We consider active sequential data release, where at each time step the user chooses from among a finite set of release mechanisms, each revealing some information about the user’s personal information, i.e., the true hypotheses, albeit with different statistics. The user manages data release in an online fashion such that maximum amount of information is revealed about the latent useful variable, while the confidence for the sensitive variable is kept below a predefined level. For the utility, we consider both the probability of correct detection of the useful variable and the mutual information (MI) between the useful variable and released data. We formulate both problems as a Markov decision process (MDP), and numerically solve them by advantage actor-critic (A2C) deep reinforcement learning (RL). Ece Naz Erdemir, Pier Luigi Dragotti, Deniz Gündüz |
ICASSP | 2 |
| 2021 | Guaranteed Reconstruction from Integrate-and-Fire Neurons with Alpha Synaptic ActivationabstractTime encoding of continuous time signals is an alternative to classical sampling paradigms. The signal is encoded in the timing of output samples rather than their amplitudes. Of particular interest are integrate-and-fire time encoding machines (IF-TEM) for sampling signals with finite rate of innovation (FRI). In contrast to state-of-the-art methods we propose an IF-TEM where we employ a biologically inspired and smooth sampling kernel, the alpha synaptic function, and show that perfect reconstruction can be achieved using this kernel. Furthermore, we derive conditions on the input signal, a train of scaled Diracs, such that not only can we guarantee the generation of useful samples, even when the Diracs have arbitrary sign, but also that these useful samples can be determined from amongst the non-useful samples. Thus, reconstruction of signals satisfying these conditions is always possible. Marek Hilton, Roxana Alexandru, Pier Luigi Dragotti |
ICASSP | 3 |
| 2021 | Model-Inspired Deep Learning for Light-Field Microscopy with Application to Neuron LocalizationabstractLight-field microscopes are able to capture spatial and angular information of incident light rays. This allows reconstructing 3D locations of neurons from a single snap-shot. In this work, we propose a model-inspired deep learning approach to perform fast and robust 3D localization of sources using light-field microscopy images. This is achieved by developing a deep network that efficiently solves a convolutional sparse coding (CSC) problem to map Epipolar Plane Images (EPI) to corresponding sparse codes. The network architecture is designed systematically by unrolling the convolutional Iterative Shrinkage and Thresholding Algorithm (ISTA) while the network parameters are learned from a training dataset. Such principled design enables the deep network to leverage both domain knowledge implied in the model, as well as new parameters learned from the data, thereby combining advantages of model-based and learning-based methods. Practical experiments on localization of mammalian neurons from light-fields show that the proposed approach simultaneously provides enhanced performance, interpretability and efficiency. Pingfan Song, Herman Verinaz-Jadan, Carmel L. Howe, Peter Quicke, Amanda J. Foust, Pier Luigi Dragotti |
ICASSP | 6 |
| 2021 | CU-Net+: Deep Fully Interpretable Network for Multi-Modal Image RestorationabstractThe network interpretability is critical in computer vision related tasks, especially for tasks involving multiple modalities. For multi-modal image restoration, one recent method, CU-Net, introduces an interpretable network based on a multi-modal convolutional sparse coding model. However, its network architecture does not mimic in full the proposed sparse model. In this paper, we overcome the limitation of CU-Net by using recurrent scheme, and this leads to a fully interpretable network which we call CU-Net+. In addition, we relax the constraint on the number of common and unique features in CU-Net, for making it more consistent with real condition. The effectiveness of the proposed CU-Net+ is evaluated on RGB guided depth image super-resolution and flash guided non-flash image denoising tasks. The numerical results show that CU-Net+ outperforms other interpretable or non-interpretable methods, with 0.16 RMSE and 0.66 dB PSNR improvement over CU-Net for the two mentioned tasks, respectively. Code is available at https://git;hub.com/JingyiXu404/CU-Net;-plus. Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
ICIP | 4 |
| 2021 | Deep Convolutional Neural Network for Multi-Modal Image Restoration and FusionabstractIn this paper, we propose a novel deep convolutional neural network to solve the general multi-modal image restoration (MIR) and multi-modal image fusion (MIF) problems. Different from other methods based on deep learning, our network architecture is designed by drawing inspirations from a new proposed multi-modal convolutional sparse coding (MCSC) model. The key feature of the proposed network is that it can automatically split the common information shared among different modalities, from the unique information that belongs to each single modality, and is therefore denoted with CU-Net, i.e., common and unique information splitting network. Specifically, the CU-Net is composed of three modules, i.e., the unique feature extraction module (UFEM), common feature preservation module (CFPM), and image reconstruction module (IRM). The architecture of each module is derived from the corresponding part in the MCSC model, which consists of several learned convolutional sparse coding (LCSC) blocks. Extensive numerical results verify the effectiveness of our method on a variety of MIR and MIF tasks, including RGB guided depth image super-resolution, flash guided non-flash image denoising, multi-focus and multi-exposure image fusion. Xin Deng 0002, Pier Luigi Dragotti |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Deep phase retrieval: Analyzing over-parameterization in phase retrieval
Junjie Huang 0001, Jubo Zhu, Wei Dai 0001, Pier Luigi Dragotti |
Signal Process. | 5 |
| 2021 | Privacy-Aware Time-Series Data Sharing With Deep Reinforcement LearningabstractInternet of things (IoT) devices are becoming increasingly popular thanks to many new services and applications they offer. However, in addition to their many benefits, they raise privacy concerns since they share fine-grained time-series user data with untrusted third parties. In this work, we study the privacy-utility trade-off (PUT) in time-series data sharing. Existing approaches to PUT mainly focus on a single data point; however, temporal correlations in time-series data introduce new challenges. Methods that preserve the privacy for the current time may leak significant amount of information at the trace level as the adversary can exploit temporal correlations in a trace. We consider sharing the distorted version of a user's true data sequence with an untrusted third party. We measure the privacy leakage by the mutual information between the user's true data sequence and shared version. We consider both the instantaneous and average distortion between the two sequences, under a given distortion measure, as the utility loss metric. To tackle the history-dependent mutual information minimization, we reformulate the problem as a Markov decision process (MDP), and solve it using asynchronous actor-critic deep reinforcement learning (RL). We evaluate the performance of the proposed solution in location trace privacy on both synthetic and GeoLife GPS trajectory datasets. For the latter, we show the validity of our solution by testing the privacy of the released location trajectory against an adversary network. Ece Naz Erdemir, Pier Luigi Dragotti, Deniz Gündüz |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Joint Learning of 3D Lesion Segmentation and Classification for Explainable COVID-19 DiagnosisabstractGiven the outbreak of COVID-19 pandemic and the shortage of medical resource, extensive deep learning models have been proposed for automatic COVID-19 diagnosis, based on 3D computed tomography (CT) scans. However, the existing models independently process the 3D lesion segmentation and disease classification, ignoring the inherent correlation between these two tasks. In this paper, we propose a joint deep learning model of 3D lesion segmentation and classification for diagnosing COVID-19, called DeepSC-COVID, as the first attempt in this direction. Specifically, we establish a large-scale CT database containing 1,805 3D CT scans with fine-grained lesion annotations, and reveal 4 findings about lesion difference between COVID-19 and community acquired pneumonia (CAP). Inspired by our findings, DeepSC-COVID is designed with 3 subnets: a cross-task feature subnet for feature extraction, a 3D lesion subnet for lesion segmentation, and a classification subnet for disease diagnosis. Besides, the task-aware loss is proposed for learning the task interaction across the 3D lesion and classification subnets. Different from all existing models for COVID-19 diagnosis, our model is interpretable with fine-grained 3D lesion distribution. Finally, extensive experimental results show that the joint learning framework in our model significantly improves the performance of 3D lesion segmentation and disease classification in both efficiency and efficacy. Xiaofei Wang 0004, Lai Jiang 0004, Liu Li 0001, Mai Xu, Xin Deng 0002, Lisong Dai, Tianyi Li 0004, Zulin Wang, Pier Luigi Dragotti |
IEEE Trans. Medical Imaging | 11 |
| 2020 | D-SLAM: Diffusion Source Localization and Trajectory MappingabstractWe consider physical fields induced by a finite number of instantaneous diffusion sources, which we sample using a mobile sensor, along unknown trajectories composed of multiple linear segments. We address the problem of estimating the sources, as well as the trajectory of the mobile sensor. Within this framework, we propose a method for localizing sources of unknown amplitudes, and known activation times. The reconstruction method we propose maps the measurements obtained using the mobile sensor to a sequence of generalized field samples. From these generalized samples, we can then retrieve the locations of the sources as well as the trajectory of the sensor (up to a linear geometric transformation). Roxana Alexandru, Thierry Blu, Pier Luigi Dragotti |
ICASSP | 3 |
| 2020 | Sampling Classes of Non-Bandlimited Signals Using Integrate-and-Fire Devices: Average Case AnalysisabstractWe investigate the use of integrate-and-fire systems to efficiently sample classes of non-bandlimited signals such as bursts of spikes. The sampling in this case is based on storing some timing information about the signal, and no information about its amplitude. We demonstrate that perfect reconstruction of these signals is possible when using proper prefiltering strategies before using a multi-channel integrate-and-fire system. We show that the probability of perfect reconstruction of the bursts of Diracs is high, even when the sufficient conditions for perfect reconstruction are not satisfied. Roxana Alexandru, Nguyen T. Thao, Dominik Rzepka, Pier Luigi Dragotti |
ICASSP | 4 |
| 2020 | Reconstruction of Fri Signals Using Deep Neural Network ApproachesabstractFinite Rate of Innovation (FRI) theory considers sampling and reconstruction of classes of non-bandlimited continuous signals that have a small number of free parameters, such as a stream of Diracs. The task of reconstructing FRI signals from discrete samples is often transformed into a spectral estimation problem and solved using Prony's method and matrix pencil method which involve estimating signal subspaces. They achieve an optimal performance given by the Cramer-Rao bound yet break down at a certain peak signal-to-́ noise ratio (PSNR). This is probably due to the so-called subspace swap event. In this paper, we aim to alleviate the subspace swap problem and investigate alternative approaches including directly estimating FRI parameters using deep neural networks and utilising deep neural networks as denoisers to reduce the noise in the samples. Simulations show significant improvements on the breakdown PSNR over existing FRI methods, which still outperform learning-based approaches in medium to high PSNR regimes. Vincent C. H. Leung, Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 3 |
| 2020 | Volume Reconstruction for Light Field MicroscopyabstractLight Field Microscopy (LFM) is a 3D imaging technique that captures volumetric information in a single snapshot. It is appealing in microscopy because of its simple implementation and the peculiarity that it is much faster than methods involving scanning. However, volume reconstruction for LFM suffers from low lateral resolution, high computational cost, and reconstruction artifacts near the native object plane. In this work, we make two contributions. First, we propose a simplification of the forward model based on a novel discretization approach that allows us to accelerate the computation without drastically increasing memory consumption. Second, we experimentally show that by including regularization priors and an appropriate initialization strategy, it is possible to remove the artifacts near the native object plane. The algorithm we use for this is ADMM. Finally, the combination of the two techniques leads to a method that outperforms classic volume reconstruction approaches (variants of Richardson-Lucy) in terms of average computational time and image quality (PSNR). Herman Verinaz-Jadan, Pingfan Song, Carmel L. Howe, Amanda J. Foust, Pier Luigi Dragotti |
ICASSP | 5 |
| 2020 | Revealing Hidden Drawings in Leonardo's 'the Virgin of the Rocks' from Macro X-Ray Fluorescence Scanning Data through Element Line LocalisationabstractMacro X-Ray Fluorescence (XRF) scanning is an increasingly widely used imaging technique for the non-invasive detection and mapping of chemical elements in Old Master paintings. Existing approaches for XRF signal analysis require varying degrees of expert user input. They are mainly based on peak fitting at fixed energies associated with each element and require the target elements to be selected manually. In this paper, we propose a new method that can process macro XRF scanning data from paintings fully automatically. The method consists of two parts: 1) detecting pulses in an XRF spectrum using Finite Rate of Innovation (FRI) theory; 2) producing the distribution maps for each element automatically identified in the painting. The results presented show the ability of our method to detect weak or partially overlapping signals and more excitingly to have visualisation of underdrawing in a masterpiece by Leonardo da Vinci. Su Yan 0003, Junjie Huang 0001, Nathan Daly, Catherine Higgitt, Pier Luigi Dragotti |
ICASSP | 5 |
| 2020 | RADAR: Robust Algorithm for Depth Image Super Resolution Based on FRI Theory and Multimodal Dictionary LearningabstractDepth image super-resolution is a challenging problem, since normally high upscaling factors are required (e.g., 16×), and depth images are often noisy. In order to achieve large upscaling factors and resilience to noise, we propose a Robust Algorithm for Depth imAge super Resolution (RADAR) that combines the power of finite rate of innovation (FRI) theory with multimodal dictionary learning. Given a low-resolution (LR) depth image, we first model its rows and columns as piece-wise polynomials and propose an FRI-based depth upscaling (FDU) algorithm to super-resolve the image. Then, the upscaled moderate quality (MQ) depth image is further enhanced with the guidance of a registered high-resolution (HR) intensity image. This is achieved by learning multimodal mappings from the joint MQ depth and HR intensity pairs to the HR depth, through a recently proposed triple dictionary learning (TDL) algorithm. Moreover, to speed up the super-resolution process, we introduce a new projection-based rapid upscaling (PRU) technique that pre-calculates the projections from the joint MQ depth and HR intensity pairs to the HR depth. Compared with the state-of-the-art deep learning-based methods, our approach has two distinct advantages: we need a fraction of training data but can achieve the best performance, and we are resilient to mismatches between training and testing datasets. The extensive numerical results show that the proposed method outperforms other state-of-the-art methods on either noise-free or noisy datasets with large upscaling factors up to 16× and can handle unknown blurring kernels well. Xin Deng 0002, Pingfan Song, Miguel R. D. Rodrigues, Pier Luigi Dragotti |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Deep Coupled ISTA Network for Multi-Modal Image Super-ResolutionabstractGiven a low-resolution (LR) image, multi-modal image super-resolution (MISR) aims to find the high-resolution (HR) version of this image with the guidance of an HR image from another modality. In this paper, we use a model-based approach to design a new deep network architecture for MISR. We first introduce a novel joint multi-modal dictionary learning (JMDL) algorithm to model cross-modality dependency. In JMDL, we simultaneously learn three dictionaries and two transform matrices to combine the modalities. Then, by unfolding the iterative shrinkage and thresholding algorithm (ISTA), we turn the JMDL model into a deep neural network, called deep coupled ISTA network. Since the network initialization plays an important role in deep network training, we further propose a layer-wise optimization algorithm (LOA) to initialize the parameters of the network before running back-propagation strategy. Specifically, we model the network initialization as a multi-layer dictionary learning problem, and solve it through convex optimization. The proposed LOA is demonstrated to effectively decrease the training loss and increase the reconstruction accuracy. Finally, we compare our method with other state-of-the-art methods in the MISR task. The numerical results show that our method consistently outperforms others both quantitatively and qualitatively at different upscaling factors for various multi-modal scenarios. Xin Deng 0002, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2019 | Time-based Sampling and Reconstruction of Non-bandlimited SignalsabstractThe last two decades have seen a renewed interest in sampling theory, which is concerned with the conversion of continuous-domain signals into discrete sequences. This conversion is traditionally achieved by recording the intensity of the signal at specified time instants. Alternatively, sampling can be based on timing rather than amplitude information. In this paper, we investigate the problem of timing-based sampling of non-bandlimited signals, within the Finite Rate of Innovation (FRI) setting. We show how these signals can be non-uniformly sampled using a compact-support kernel that satisfies the generalised Strang-Fix conditions, and a comparator. We then prove that perfect input estimation is possible using a novel local reconstruction algorithm. Roxana Alexandru, Pier Luigi Dragotti |
ICASSP | 2 |
| 2019 | Coupled Ista Network for Multi-modal Image Super-resolutionabstractIn this paper, we propose a novel deep neural network architecture for multi-modal image super-resolution (MISR). The architecture is based on a new joint multi-modal dictionary learning (JMDL) algorithm to model cross-modality dependency and to map them to a high-resolution version of one modality. In JMDL, we learn three dictionaries and two transform matrices to combine the modalities. By using the learned model, we then design the network architecture by a coupled unfolding of the iterative shrinkage and thresholding algorithm (ISTA). We finally initialize the parameters of our network with a new optimization strategy. The initialized parameters are demonstrated to effectively decrease the training loss and increase the reconstruction accuracy. The numerical results show that our method outperforms other state-of-the-art methods quantitatively and qualitatively for MISR. Xin Deng 0002, Pier Luigi Dragotti |
ICASSP | 2 |
| 2019 | Privacy-cost Trade-off in a Smart Meter System with a Renewable Energy Source and a Rechargeable BatteryabstractWe study the privacy-cost trade-off in a smart meter (SM) system with a renewable energy source (RES) and a finite-capacity rechargeable battery (RB). Privacy is measured by the mutual information rate between the energy demand and the energy received from the grid, where the latter also determines the cost, and hence, reported by the SM to the utility provider (UP). We consider a renewable energy generation process that fully charges the RB at random time instants, and its realization is assumed to be known also by the UP. We reformulate the problem as a Markov decision process (MDP), and solve it by dynamic programming (DP) to design battery charging and discharging policies that minimize a linear combination of the privacy leakage and energy cost. We also propose a lower bound and two alternative low-complexity energy management policies, one of which is shown numerically to perform close to the MDP solution. Ece Naz Erdemir, Pier Luigi Dragotti, Deniz Gündüz |
ICASSP | 2 |
| 2019 | A Deep Dictionary Model to Preserve and Disentangle Key Features in a SignalabstractWe propose a deep dictionary model for single image super-resolution (SISR) made of multiple layers of analysis dictionaries interlaced with corresponding soft-thresholding operations and a single synthesis dictionary. In this paper, we introduce a novel method for learning analysis dictionary and thresholding pairs as building block for the deep dictionary model. Each analysis dictionary contains two sub-dictionaries: an information preserving analysis dictionary (IPAD) and a clustering analysis dictionary (CAD). The IPAD and thresholding pair passes the key information from the previous layer, while the CAD and thresholding pair gives a sparse representation of its input data that facilitates discrimination of key features. Simulation results show that the proposed deep dictionary model achieves comparable performance with a deep neural network which has the same structure and is optimized using backpropagation. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 2 |
| 2019 | Wavelet Domain Style Transfer for an Effective Perception-Distortion Tradeoff in Single Image Super-ResolutionabstractIn single image super-resolution (SISR), given a low-resolution (LR) image, one wishes to find a high-resolution (HR) version of it which is both accurate and photorealistic. Recently, it has been shown that there exists a fundamental tradeoff between low distortion and high perceptual quality, and the generative adversarial network (GAN) is demonstrated to approach the perception-distortion (PD) bound effectively. In this paper, we propose a novel method based on wavelet domain style transfer (WDST), which achieves a better PD tradeoff than the GAN based methods. Specifically, we propose to use 2D stationary wavelet transform (SWT) to decompose one image into low-frequency and high-frequency sub-bands. For the low-frequency sub-band, we improve its objective quality through an enhancement network. For the high-frequency sub-band, we propose to use WDST to effectively improve its perceptual quality. By feat of the perfect reconstruction property of wavelets, these sub-bands can be re-combined to obtain an image which has simultaneously high objective and perceptual quality. The numerical results on various datasets show that our method achieves the best trade-off between the distortion and perceptual quality among the existing state-of-the-art SISR methods. Xin Deng 0002, Mai Xu, Pier Luigi Dragotti |
ICCV | 4 |
| 2018 | U-Fresh: An Fri-Based Single Image Super Resolution Algorithm and An Application in Image CompressionabstractLearning based single image super resolution (SISR) methods have achieved notable results, however, they require large datasets for training, and may struggle when there is a mismatch between the testing and training data. To overcome these drawbacks, we propose an approach, named U - FRESH, which only requires a small dataset but can achieve state-of-the-art performance also in the presence of training and testing mismatches. We accomplish this by leveraging a method called FRESH, which enhances the image resolution using FRI theory. We start upscaling from the FRESH generated low resolution image. To minimize the reconstruction error, we propose a new regression selection technique to make the mapping more reliable and robust, and a wavelet based back projection technique to improve the quality of the reconstructed image. Based on U - FRESH, we also propose a new framework based on JPEG 2000 for image compression. Numerical results show that our U-FRESH method achieves state-of-the-art performance in SISR and provides better compression results than JPEG 2000. Xin Deng 0002, Junjie Huang 0001, Mengying Liu, Pier Luigi Dragotti |
ICASSP | 4 |
| 2018 | A Deep Dictionary Model for Image Super-ResolutionabstractInspired by the recent success of deep neural network architectures and the recent effort to develop multi-layer sparse models, we propose a novel deep dictionary learning architecture which is optimized to address a specific regression task known as single image super-resolution. Contrary to other multi-layer dictionaries, our architecture contains L-1 analysis dictionaries to extract high-level features and one synthesis dictionary which is designed to optimize the regression task. We propose a variation of an existing method to learn the analysis dictionaries and we update them without the need to use a back-propagation approach. Results on image super-resolution are satisfactory. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 2 |
| 2018 | CosMIC: A Consistent Metric for Spike Inference from Calcium ImagingabstractIn recent years, the development of algorithms to detect neuronal spiking activity from two-photon calcium imaging data has received much attention, yet few researchers have examined the metrics used to assess the similarity of detected spike trains with the ground truth. We highlight the limitations of the two most commonly used metrics, the spike train correlation and success rate, and propose an alternative, which we refer to as CosMIC. Rather than operating on the true and estimated spike trains directly, the proposed metric assesses the similarity of the pulse trains obtained from convolution of the spike trains with a smoothing pulse. The pulse width, which is derived from the statistics of the imaging data, reflects the temporal tolerance of the metric. The final metric score is the size of the commonalities of the pulse trains as a fraction of their average size. Viewed through the lens of set theory, CosMIC resembles a continuous Sørensen-Dice coefficient-an index commonly used to assess the similarity of discrete, presence/absence data. We demonstrate the ability of the proposed metric to discriminate the precision and recall of spike train estimates. Unlike the spike train correlation, which appears to reward overestimation, the proposed metric score is maximized when the correct number of spikes have been detected. Furthermore, we show that CosMIC is more sensitive to the temporal precision of estimates than the success rate. Stephanie Reynolds, Therese Abrahamsson, Jesper Sjöström, Simon R. Schultz, Pier Luigi Dragotti |
Neural Comput. | 5 |
| 2018 | Photo Realistic Image Completion via Dense CorrespondenceabstractIn this paper, we propose an image completion algorithm based on dense correspondence between the input image and an exemplar image retrieved from the Internet. Contrary to traditional methods which register two images according to sparse correspondence, in this paper, we propose a hierarchical PatchMatch method that progressively estimates a dense correspondence, which is able to capture small deformations between images. The estimated dense correspondence has usually large occlusion areas that correspond to the regions to be completed. A nearest neighbor field (NNF) interpolation algorithm interpolates a smooth and accurate NNF over the occluded region. Given the calculated NNF, the correct image content from the exemplar image is transferred to the input image. Finally, as there could be a color difference between the completed content and the input image, a color correction algorithm is applied to remove the visual artifacts. Numerical results show that our proposed image completion method can achieve photo realistic image completion results. Junjie Huang 0001, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2018 | Sparse Representation in Fourier and Local Bases Using ProSparse: A Probabilistic AnalysisabstractFinding the sparse representation of a signal in an overcomplete dictionary has attracted a lot of attention over the past years. This paper studies ProSparse, a new polynomial complexity algorithm that solves the sparse representation problem when the underlying dictionary is the union of a Vandermonde matrix and a banded matrix. Unlike our previous work, which establishes deterministic (worst-case) sparsity bounds for ProSparse to succeed, this paper presents a probabilistic average-case analysis of the algorithm. Based on a generating-function approach, closed-form expressions for the exact success probabilities of ProSparse are given. The success probabilities are also analyzed in the high-dimensional regime. This asymptotic analysis characterizes a sharp phase transition phenomenon regarding the performance of the algorithm. Yue M. Lu, Jon Onativia, Pier Luigi Dragotti |
IEEE Trans. Inf. Theory | 3 |
| 2018 | DAGAN: Deep De-Aliasing Generative Adversarial Networks for Fast Compressed Sensing MRI ReconstructionabstractCompressed sensing magnetic resonance imaging (CS-MRI) enables fast acquisition, which is highly desirable for numerous clinical applications. This can not only reduce the scanning cost and ease patient burden, but also potentially reduce motion artefacts and the effect of contrast washout, thus yielding better image quality. Different from parallel imaging-based fast MRI, which utilizes multiple coils to simultaneously receive MR signals, CS-MRI breaks the Nyquist-Shannon sampling barrier to reconstruct MRI images with much less required raw data. This paper provides a deep learning-based strategy for reconstruction of CS-MRI, and bridges a substantial gap between conventional non-learning methods working only on data from a single image, and prior knowledge from large training data sets. In particular, a novel conditional Generative Adversarial Networks-based model (DAGAN)-based model is proposed to reconstruct CS-MRI. In our DAGAN architecture, we have designed a refinement learning method to stabilize our U-Net based generator, which provides an end-to-end network to reduce aliasing artefacts. To better preserve texture and edges in the reconstruction, we have coupled the adversarial loss with an innovative content loss. In addition, we incorporate frequency-domain information to enforce similarity in both the image and frequency domains. We have performed comprehensive comparison studies with both conventional CS-MRI reconstruction methods and newly investigated deep learning approaches. Compared with these methods, our DAGAN method provides superior reconstruction with preserved perceptual image details. Furthermore, each image is reconstructed in about 5 ms, which is suitable for real-time processing. Guang Yang 0006, Simiao Yu, Hao Dong 0003, Gregory Slabaugh, Pier Luigi Dragotti, Xujiong Ye, Fangde Liu, Simon R. Arridge, Jennifer Keegan, Yike Guo, David N. Firmin |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Solving Inverse Source Problems for Sources with Arbitrary Shapes using Sensor Networks
John Murray-Bruce, Pier Luigi Dragotti |
ESANN | 2 |
| 2017 | ProSparse extension: Prony's based sparse pattern recovery with extended dictionariesabstractProSparse is a Prony's based method that solves the sparse representation problem of signals in the union of Fourier and canonical bases. By exploiting the structure of the dictionary, ProSparse is able to reconstruct sparse signals beyond the recovery bound of Basis Pursuit. We generalize this framework for a broader class of dictionaries which are still formed from the union of two bases. The proposed algorithm achieves perfect reconstruction over a lower sparsity level than Basis Pursuit in noiseless cases. In the presence of noise, we extend the ProSparse Denoise algorithm to the generalized dictionaries by considering their intrinsic structure. The original ProSparse can be viewed as a special case of our proposed algorithm. From simulation results, our approach outperforms state-of-the-art algorithms. Junjie Huang 0001, Pier Luigi Dragotti |
ICASSP | 2 |
| 2017 | Identifying a multiple plane plenoptic function from a swiped imageabstractBlur in images, caused by camera motion with an open shutter, is usually thought of as a problem. The algorithm described in this paper shows instead that it is possible to use the blur caused by the integration of light rays at different locations along a moving camera trajectory to extract information about the light rays that are present within the scene. Retrieving the light rays present within a scene from different viewpoints is equivalent to retrieving the plenoptic function of the scene. In this paper, we focus on a specific case in which the blurred image of a scene, containing fronto-parallel planes with uniform unknown textures, is analysed to recreate the plenoptic function. The image is captured by a digital single lens camera with shutter open, moving in a straight line between two points, resulting in a swiped image. We estimate the EPI from this blurred image, and the EPI can be used to generate unblurred images for a given camera location. Michael Lawson, Mike Brookes, Pier Luigi Dragotti |
ICASSP | 3 |
| 2017 | Model order selection for sampling FRI signalsabstractRecently it has been shown that specific classes of non-bandlimited signals known as signals with finite rate of innovation (FRI) can be perfectly reconstructed by using appropriate sampling kernels and reconstruction schemes. The knowledge of the model order (i.e. the rate of innovation) is essential for correct reconstruction. In view of this, we devise an algorithm which can robustly identify the rate of innovation prior to the signal reconstruction in different noise levels and this extends the current scheme to a universal one that works with signals with unknown rate of innovation and using arbitrary kernels. We use the `guaranteed performance' criterion to assess the performance and show a success rate close to 100% for SNR up to 10dB. Xiaoyao Wei, Pier Luigi Dragotti |
ICASSP | 2 |
| 2016 | The graph FRI framework-spline wavelet theory and sampling on circulant graphsabstractThe objective of this work is to consider sparse representations of certain classes of signals on circulant graphs, by introducing families of graph wavelets which possess vanishing (exponential) moment properties. In light of this, we propose a novel framework of sampling and perfect reconstruction of sparse and wavelet-sparse signals on circulant graphs, which we denote as the Graph FRI framework, as an extension to the traditional discrete case. Given the dimensionality-reduced GFT of a sparse signal on a graph G, we can perfectly reconstruct the latter, while inferring a distinct down-sampling pattern and the structure of the associated coarsened graph through decomposition of the GFT-basis as the product between a coefficient matrix C and the multiresolution filtering operation with a low-pass graph e-spline filter. Hereby, we demonstrate that for a sufficiently banded adjacency matrix A of G, the obtained coarse graph preserves the original generating set S of G in a scheme of spectral sampling with respect to the original eigenbasis of A. Madeleine S. Kotzagiannidis, Pier Luigi Dragotti |
ICASSP | 2 |
| 2016 | Reconstructing non-point sources of diffusion fields using sensor measurementsabstractWe present a framework for estimating non-localized sources of diffusion fields using spatiotemporal measurements of the field. Specifically in this contribution, we consider two non-localized source types: straight line and polygonal sources and assume that the induced field is monitored using a sensor network. Given the sensor measurements, we demonstrate, for each non-point source parameterization, how to reduce the source estimation problem to a system governed by a power series expansion that can then be efficiently solved using Prony's method, in order to reconstruct the source. We then evaluate the proposed algorithms by performing some numerical simulations using both noiseless and noisy spatiotemporal sensor measurements of the field. John Murray-Bruce, Pier Luigi Dragotti |
ICASSP | 2 |
| 2016 | Prosparse denoise: Prony's based sparse pattern recovery in the presence of noiseabstractWe present a novel algorithm - ProSparse Denoise - that can solve the sparsity recovery problem in the presence of noise when the dictionary is the union of Fourier and identity matrices. The algorithm is based on a proper use of Cadzow routine and Prony's method and exploits the duality of Fourier and identity matrices. The algorithm has low complexity compared to state of the art algorithms for sparse recovery since it relies on the Fast Fourier Transform (FFT) algorithm. We provide conditions on the noise that guarantees the correct recovery of the sparsity pattern. Our approach outperforms state of the art algorithms such as Basis Pursuit De-noise and Subspace Pursuit when the dictionary is the union of Fourier and identity matrices. Jon Onativia, Yue M. Lu, Pier Luigi Dragotti |
ICASSP | 3 |
| 2016 | Video temporal super-resolution using nonlocal registration and self-similarityabstractIn this paper we present a novel temporal super-resolution method for increasing the frame-rate of single videos. The proposed algorithm is based on motion-compensated 3-D patches, i.e., a sequence of 2-D blocks following a given motion trajectory. The trajectories are computed through a coarse-to-fine motion estimation strategy embedding a regularized block-wise distance metric that takes into account the coherence of neighbouring motion vectors. Our algorithm comprises two stages. In the first stage, a nonlocal search procedure is used to find a set of 3-D patches (targets) similar to a given patch (reference), subsequently all targets are registered at sub-pixel precision with respect to the reference in an upsampled 3-D FFT domain, and finally all registered patches are aggregated at their appropriate locations in the high-resolution video. The second stage is used to further improve the estimation quality by correcting each 3-D patch of the video obtained from the first stage with a linear operator learned from the self-similarity of patches at a lower temporal scale. Our experimental evaluation on color videos shows that the proposed approach achieves high quality super-resolution results from both an objective and subjective point of view. Matteo Maggioni, Pier Luigi Dragotti |
MMSP | 2 |
| 2016 | Identification of Transform Coding ChainsabstractTransform coding is routinely used for lossy compression of discrete sources with memory. The input signal is divided into N-dimensional vectors, which are transformed by means of a linear mapping. Then, transform coefficients are quantized and entropy coded. In this paper, we consider the problem of identifying the transform matrix as well as the quantization step sizes. First, we study the case in which the only available information is a set of P transform decoded vectors. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors. We propose an algorithm that is able to find the optimal solution and we formally study its convergence properties. Three potential realms of application are considered as example scenarios for the proposed theory: 1) parameter retrieval in the presence of a chain of two transform coders; 2) image tampering identification; and 3) parameter estimation for predictive coders. We show that, despite their differences, all three scenarios can be tackled by applying the same fundamental methodology. Experiments on both the synthetic data and the real images validate the proposed approach. Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro |
IEEE Trans. Image Process. | 3 |
| 2016 | FRESH - FRI-Based Single-Image Super-Resolution AlgorithmabstractIn this paper, we consider the problem of single image super-resolution and propose a novel algorithm that outperforms state-of-the-art methods without the need of learning patches pairs from external data sets. We achieve this by modeling images and, more precisely, lines of images as piecewise smooth functions and propose a resolution enhancement method for this type of functions. The method makes use of the theory of sampling signals with finite rate of innovation (FRI) and combines it with traditional linear reconstruction methods. We combine the two reconstructions by leveraging from the multi-resolution analysis in wavelet theory and show how an FRI reconstruction and a linear reconstruction can be fused using filter banks. We then apply this method along vertical, horizontal, and diagonal directions in an image to obtain a single-image super-resolution algorithm. We also propose a further improvement of the method based on learning from the errors of our super-resolution result at lower resolution levels. Simulation results show that our method outperforms state-of-the-art algorithms under different blurring kernels. Xiaoyao Wei, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2015 | Consensus for the distributed estimation of point diffusion sources in sensor networksabstractIn this contribution, we implement a fully distributed diffusion field estimation algorithm based on the use of average consensus schemes. We show that the field reconstruction problem is equivalent to estimating the sources of the field, and then derive an exact inversion formula for jointly recovering these sources when they are localized and instantaneous. Next we adapt this formula to the sensor network setting when only spatiotemporal samples of the field are available, and only local interactions between the sensors are allowed. To this end, we propose a robust distributed algorithm for reconstructing two-dimensional diffusion fields, sampled with a network of arbitrarily placed sensors. The proposed distributed algorithm is validated through numerical simulations in the noisy, multiple source setting. John Murray-Bruce, Pier Luigi Dragotti |
ICASSP | 2 |
| 2015 | Sparsity pattern recovery using FRI methodsabstractThe problem of finding the sparse representation of a signal has attracted a lot of attention over the past years. In particular, uniqueness conditions and reconstruction algorithms have been established by relaxing a non-convex optimisation problem. The finite rate of innovation (FRI) theory is an alternative approach that solves the sparsity problem using algebraic methods based around Prony's algorithm. Recent extensions to this framework have shown that it is possible to recover sparse representations beyond the uniqueness limits, that is, finding all the possible sparse representations that fit the observation for the case of signals which are sparse in the union of Fourier and canonical bases. In this paper, we show the application of such methods to the case of the union of DCT and Haar basis. We present an extension that takes advantage of the even symmetry of the cosine functions to build an algorithm that can operate over the observed vector and in a dual domain. We also analyse the case of the union of frames. Simulation results confirm the validity of this new approach and show that it outperforms state of the art algorithms in a number scenarios. Jon Onativia, Yue M. Lu, Pier Luigi Dragotti |
ICASSP | 3 |
| 2015 | Sampling piecewise smooth signals and its application to image up-samplingabstractIn this paper we consider the problem of sampling piecewise smooth signals. The classical sampling theory is not able to sample them efficiently since the signals are neither bandlimited nor live in a shift-invariant subspace. We propose a sampling scheme which still use the set of samples from a classical sampling set-up with an arbitrary acquisition device but achieves accurate reconstruction by making use of the non-linear approximate FRI reconstruction method in addition to the classical linear reconstruction. Specifically, we see the class of piecewise smooth signal as the sum of a piecewise polynomial, which is a signal that is fully specified by finite number of parameters and can be recovered with FRI non-linear methods, and a globally smooth term, which can be accurately recovered by linear reconstruction. We also show that our proposed scheme based on combining linear and non-linear reconstruction methods can be employed for resolution enhancement of discrete-time piece-wise smooth signals and images. Xiaoyao Wei, Pier Luigi Dragotti |
ICIP | 2 |
| 2015 | Guaranteed Performance in the FRI SettingabstractFinite Rate of Innovation (FRI) sampling theory has shown that it is possible to sample and perfectly reconstruct classes of non-bandlimited signals such as streams of Diracs. In the case of noisy measurements, FRI methods achieve the optimal performance given by the Cramér-Rao bound up to a certain PSNR and breaks down for smaller PSNRs. To the best of our knowledge, the precise anticipation of the breakdown event in FRI settings is still an open problem. In this letter, we address this issue by investigating the subspace swap event which has been broadly recognised as the reason for performance breakdown in SVD-based parameter estimation algorithms. We work out at which noise level the absence of subspace swap is guaranteed and this gives us an accurate prediction of the breakdown PSNR which we also relate to the sampling rate and the distance between adjacent Diracs. Simulation results validate the reliability of our analysis. Xiaoyao Wei, Pier Luigi Dragotti |
IEEE Signal Process. Lett. | 2 |
| 2015 | On the Reconstruction of Wavelet-Sparse Signals From Partial Fourier InformationabstractThe problem of reconstructing a wavelet-sparse signal from its partial Fourier information has received a lot of attention since the emergence of compressive sensing (CS). The latest theory within the CS framework analyzes the local coherence between the Fourier and wavelet bases, and recover the signal from frequencies randomly selected according to a variable density profile. Unlike these developments, we adopt a new approach that does not need to analyze the (local) coherence. We show that the problem can be tackled by recovering the wavelet coefficients from the finest to the coarse scale, and only a small set of frequencies are needed to recover the coefficients exactly. As long as the scaling function satisfies a mild condition, the reconstruction is exact. Moreover the frequency set can be deterministically pre-selected and does not need to change even if the wavelet basis changes. Yingsong Zhang, Pier Luigi Dragotti |
IEEE Signal Process. Lett. | 2 |
| 2015 | An Image Recapture Detection Algorithm Based on Learning Dictionaries of Edge ProfilesabstractWith today's digital camera technology, high-quality images can be recaptured from an liquid crystal display (LCD) monitor screen with relative ease. An attacker may choose to recapture a forged image in order to conceal imperfections and to increase its authenticity. In this paper, we address the problem of detecting images recaptured from LCD monitors. We provide a comprehensive overview of the traces found in recaptured images, and we argue that aliasing and blurriness are the least scene dependent features. We then show how aliasing can be eliminated by setting the capture parameters to predetermined values. Driven by this finding, we propose a recapture detection algorithm based on learned edge blurriness. Two sets of dictionaries are trained using the K-singular value decomposition approach from the line spread profiles of selected edges from single captured and recaptured images. An support vector machine classifier is then built using dictionary approximation errors and the mean edge spread width from the training images. The algorithm, which requires no user intervention, was tested on a database that included more than 2500 high-quality recaptured images. Our results show that our method achieves a performance rate that exceeds 99% for recaptured images and 94% for single captured images. Thirapiroon Thongkamwitoon, Hani Muammar, Pier Luigi Dragotti |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Spatio-temporal sampling and reconstruction of diffusion fields induced by point sourcesabstractIn this paper we consider a diffusion field induced by multiple point sources and address the problem of reconstructing the field from its spatio-temporal samples obtained using a sensor network. We begin by formulating the problem as a multi-source estimation problem - so estimating source locations, activation times and intensities given samples of the induced field. Next a two-step algorithm is proposed for the single (localized and instantaneous) source field. First, the source location and intensity are estimated by applying the “reciprocity gap” principle; we show that this step can also reveal locations of multiple non-instantaneous sources. In the second step, we use an iterative method, based on Cauchy-Schwarz inequality, to find the activation time given the estimated location and intensity. Finally we extend this algorithm to the multi-source field and present simulation results to validate our findings. John Murray-Bruce, Pier Luigi Dragotti |
ICASSP | 2 |
| 2014 | Universal sampling of signals with finite rate of innovationabstractRecently it has been shown that specific classes of non-bandlimited signals known as signals with finite rate of innovation (FRI) can be perfectly reconstructed by using appropriate sampling kernels and reconstruction schemes. This exact FRI framework was later extended to an approximate FRI framework that works with any kernel. Reconstruction is achieved by recovering all the parameters in the parametric model of the incoming signal, hence it is essential to know the model order (the rate of innovation) to ensure recovery. In view of this, we devise an algorithm for identifying the rate of innovation in order to extend the current sampling scheme to a universal one which enables sampling signals with arbitrary FRI using any acquisition device. Our proposed algorithm can effectively identify the rate of innovation prior to the signal reconstruction using arbitrary kernels and in different noise levels where we also show that it achieves the performance predicted by the Cramèr-Rao bounds. Xiaoyao Wei, Pier Luigi Dragotti |
ICASSP | 2 |
| 2014 | The modulated E-spline with multiple subbands and its application to sampling wavelet-sparse signalsabstractThe theory of Finite Rate of Innovation (FRI) can be applied to sampling and reconstructing certain classes of parametric signals. The objective of this paper is to have a sub-Nyquist sampling scheme for continuous-time wavelet-sparse signals within the general framework of FRI theory. Though the signal has a parametric representation in the wavelet basis, it is not possible to recover the signal merely from its low-pass samples, which makes the problem different from the conventional FRI settings. The need for the Fourier coefficients at frequencies widely spread over the spectrum puts challenges on the design of the sampling kernel. This paper presents a new family of sampling kernels that are able to stably reproduce exponentials over a wide range of frequencies and gives numerical examples on applying the new kernel to sampling wavelet-sparse signals. Yingsong Zhang, Pier Luigi Dragotti |
ICASSP | 2 |
| 2014 | Wide-baseline image change detectionabstractWe present a fully automated method for the detection of changes within a scene between a reference and a sample image whose viewing angles differ by up to 30°. We also describe an extension to the SIFT technique that allows extracted feature points to be matched over wider viewing angles. Matched correspondences between reference and sample images are used to construct a Delaunay triangulation and changes are detected by comparing triangles after affine compensation using a dense SIFT metric. False positives are reduced by using a novel technique introduced as local plane matching (LPM) to match mean-shift segments in unmatched areas using the homographies of local planes to compensate for perspective distortions. The method is shown to achieve pixel-level equal error rates of 5% at a 10° azimuth view angle difference. Ziggy Jones, Mike Brookes, Pier Luigi Dragotti, David M. Benton |
ICIP | 3 |
| 2014 | Tilted layer-based modeling for enhanced light-field processing and image based renderingabstractImage based rendering is an attractive approach for novel view synthesis due to its low complexity requirements and potential for photorealistic results. However for successful rendering, geometric priors about the structure of the scene are necessary. In this paper we present a tilted layer model approximation of the plenoptic function which gives improved modeling of scenes where the objects are not fronto-parallel to the camera views while preserving occlusion ordering. The framework is extended to the case where camera positions are not constrained to a single plane but can lie on multiple planes. Results on the Middlebury dataset and simulated scenes show that better rendering results can be obtained compared with the state-of-the-art using a fronto-parallel layer model, or alternatively similar results can be obtained with a more compact layer representation of the scene. James Pearson, Marco Visentini Scarzanella, Mike Brookes, Pier Luigi Dragotti |
ICIP | 4 |
| 2014 | Robust image recapture detection using a K-SVD learning approach to train dictionaries of edge profilesabstractA professionally recaptured image from an LCD monitor can be, visually, very difficult to distinguish from its original counterpart. In this paper we show that it is possible to detect a recaptured image from the unique nature of the edge profiles present in the image. We leverage the fact that the edge profiles of single and recaptured images are markedly different and we train two alternative dictionaries using the K-SVD approach. One dictionary is trained to provide a sparse representation of single captured edges and a second for recaptured edges. Using these two learned dictionaries, we can determine whether a query image has been recaptured. We achieve this by observing the type of dictionary that gives the smallest error in a sparse representation of the edges of the query image. Experiments conducted show that the proposed algorithm is capable of detecting recaptured images with a high level of accuracy and copes well with a wide range of natural images. Thirapiroon Thongkamwitoon, Hani Muammar, Pier Luigi Dragotti |
ICIP | 3 |
| 2014 | On the Spectrum of the Plenoptic FunctionabstractThe plenoptic function is a powerful tool to analyze the properties of multi-view image data sets. In particular, the understanding of the spectral properties of the plenoptic function is essential in many computer vision applications, including image-based rendering. In this paper, we derive for the first time an exact closed-form expression of the plenoptic spectrum of a slanted plane with finite width and use this expression as the elementary building block to derive the plenoptic spectrum of more sophisticated scenes. This is achieved by approximating the geometry of the scene with a set of slanted planes and evaluating the closed-form expression for each plane in the set. We then use this closed-form expression to revisit uniform plenoptic sampling. In this context, we derive a new Nyquist rate for the plenoptic sampling of a slanted plane and a new reconstruction filter. Through numerical simulations, on both real and synthetic scenes, we show that the new filter outperforms alternative existing filters. Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes |
IEEE Trans. Image Process. | 2 |
| 2014 | Quadtree Structured Image Approximation for Denoising and InterpolationabstractThe success of many image restoration algorithms is often due to their ability to sparsely describe the original signal. Shukla proposed a compression algorithm, based on a sparse quadtree decomposition model, which could optimally represent piecewise polynomial images. In this paper, we adapt this model to the image restoration by changing the rate-distortion penalty to a description-length penalty. In addition, one of the major drawbacks of this type of approximation is the computational complexity required to find a suitable subspace for each node of the quadtree. We address this issue by searching for a suitable subspace much more efficiently using the mathematics of updating matrix factorisations. Algorithms are developed to tackle denoising and interpolation. Simulation results indicate that we beat state of the art results when the original signal is in the model (e.g., depth images) and are competitive for natural images when the degradation is high. Adam Scholefield, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2014 | On Sparse Representation in Fourier and Local BasesabstractWe consider the classical problem of finding the sparse representation of a signal in a pair of bases. When both bases are orthogonal, it is known that the sparse representation is unique when the sparsity K of the signal satisfies K <; 1/μ(D), where μ(D) is the mutual coherence of the dictionary. Furthermore, the sparse representation can be obtained in polynomial time by basis pursuit (BP), when K <; 0.91/μ(D). Therefore, there is a gap between the unicity condition and the one required to use the polynomial-complexity BP formulation. For the case of general dictionaries, it is also well known that finding the sparse representation under the only constraint of unicity is NP-hard. In this paper, we introduce, for the case of Fourier and canonical bases, a polynomial complexity algorithm that finds all the possible K-sparse representations of a signal under the weaker condition that K <; √2/μ(D). Consequently, when K <; 1/μ(D), the proposed algorithm solves the unique sparse representation problem for this structured dictionary in polynomial time. We further show that the same method can be extended to many other pairs of bases, one of which must have local atoms. Examples include the union of Fourier and local Fourier bases, the union of discrete cosine transform and canonical bases, and the union of random Gaussian and canonical bases. Pier Luigi Dragotti, Yue M. Lu |
IEEE Trans. Inf. Theory | 1 |
| 2013 | Quantisation Invariants for Transform Parameter Estimation in Coding ChainsabstractWe examine the case of a signal going through a processing chain consisting of two transform coding stages, with the aim of recovering the unknown parameters of the first encoder. Through number theoretical considerations, we identify a lattice of quantisation invariant points, whose coordinates are not affected by the double quantisation and whose parameters are closely related to the unknown transform. The conditions for this lattice to exist are then discussed, and its uniqueness properties analysed. Finally, an algorithmic procedure to recover the invariants from a sparse set of points is shown together with numerical results. Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti |
DCC | 3 |
| 2013 | An investigation into aliasing in images recaptured from an LCD monitor using a digital cameraabstractWith current technology, high quality recaptured images can be created from soft displays, such as an LCD monitor, using a digital still camera and professional image editing software. The task of verifying the ownership and past history of an image is, consequently, more difficult. One approach to detecting an image that has been recaptured from an LCD monitor is to search for the presence of aliasing due to the sampling of the monitor pixel grid. To validate this approach, an investigation into the aliasing introduced in a digitally recaptured image is conducted. An anti-forensic method for recapturing images that are free from aliasing is developed using a model of the image acquisition process. This is supported by a simulation of the acquisition process and illustrated with examples of recaptured images that are free from aliasing. Hani Muammar, Pier Luigi Dragotti |
ICASSP | 2 |
| 2013 | Sequential local FRI sampling of infinite streams of DiracsabstractThe theory of sampling signals with finite rate of innovation (FRI) has shown that it is possible to perfectly recover classes of non-bandlimited signals such as streams of Diracs from uniform samples. Most of previous papers, however, have to some extent only focused on the sampling of periodic or finite duration signals. In this paper we propose a novel method that is able to reconstruct infinite streams of Diracs, even in high noise scenarios. We sequentially process the discrete samples and output locations and amplitudes of the Diracs in real-time. We first establish conditions for perfect reconstruction in the noiseless case and then present the sequential algorithm for the noisy scenario. We also show that we can achieve a high reconstruction accuracy of 1000 Diracs for SNRs as low as 5dB. Jon Onativia, Jose Antonio Uriguen, Pier Luigi Dragotti |
ICASSP | 3 |
| 2013 | Transform coder identificationabstractThe widespread popularity of transform coding has made it central to a wide range of methods in forensics, quality assessment and digital restoration. However, most approaches require prior knowledge of the transform coding parameters. In this paper, we consider the challenging problem of identifying the transform matrix as well as the quantization step sizes of a transform coder, given a set of P non-overlapping N-dimensional vectors observed as its output. We formulate the problem in terms of finding the lattice with the largest determinant that contains all observed vectors and we propose an algorithm that is able to find the optimal solution. Our experimental analysis shows that the probability of success of the algorithm quickly approaches 1 for small values of (P - N). The complexity of the proposed algorithm grows linearly with the dimensionality N. Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro |
ICASSP | 3 |
| 2013 | Reverse engineering of signal acquisition chains using the theory of sampling signals with finite rate of innovationabstractThis paper presents a novel theoretical framework for the reverse engineering of signal acquisition chains. We investigate how signals are transformed through a chain of signal acquisition and reconstruction stages. The signals, at different stages in the chain, are modelled using the theory of sampling signals with Finite Rate of Innovation (FRI). The model allows us to determine the chain structure and corresponding acquisition history from unknown query signals and to retrieve important parameters relating to the acquisition chain. Thirapiroon Thongkamwitoon, Hani Muammar, Pier Luigi Dragotti |
ICASSP | 3 |
| 2013 | Video recapture detection based on ghosting artifact analysisabstractVideo forensics is becoming a popular field of research and an increasing number of forensic techniques have been proposed in the last few years. However, a simple yet effective method to fool many detectors consists in recapturing a video sequence with a camcorder. For this reason being able to detect video recapture is a topic of interest for a forensic analyst. In this paper, we first characterize the video recapture model, focusing on the common scenario of a sequence recaptured from a LCD monitor using a digital camcorder, then we propose a recapture detector for this case. The detector is based on the analysis of a characteristic ghosting artifact left by the recapture process. The presented algorithm is finally validated by means of tests on original and recaptured sequences. These tests prove that the algorithm achieves high accuracy results. Paolo Bestagini, Marco Visentini Scarzanella, Marco Tagliasacchi, Pier Luigi Dragotti, Stefano Tubaro |
ICIP | 4 |
| 2013 | Transform coder identification with double quantized dataabstractThe analysis of chains of double transform coders has been recently addressed in the image forensic literature, especially for the case of double JPEG compression. In that case, the transform is assumed to be known a priori (e.g., 2D-DCT), whereas the quantization steps of the first coder need to be determined. In this work, we generalize the analysis to the challenging case in which nothing is known about the first coder, but that the transform is orthonormal. Given a set of vectors observed as output of a chain of two transform coders, we identify both the transform and the quantizer of the first. The key idea is to denoise the observed vectors exploiting the constraints imposed by the first quantizer and then apply our previously proposed method, which successfully performs transform identification in the case of noiseless observations. Experiments on real images validate the proposed approach. Marco Tagliasacchi, Marco Visentini Scarzanella, Pier Luigi Dragotti, Stefano Tubaro |
ICIP | 3 |
| 2013 | Modelling radial distortion chains for video recapture detectionabstractThis paper presents a novel cue for automatic recapture detection of videos. The problem of recapture detection is important to the field of digital forensics as recapture is often an indicator of prior tampering activity. In this paper, we tackle the problem by considering the deformation underwent by geometric primitives, such as straight lines, when processed along recapture chains.We mathematically derive a general curve model for straight lines deformed after single capture under a radial distortion model. The model is then extended to the case of recapture, demonstrating how to automatically classify videos on a per-frame basis from its compliance with a low-order radial distortion model. Finally, we test our model with a practical detector to automatically extract deformed straight lines for classification, which is applied to synthetic sequences. Marco Visentini Scarzanella, Pier Luigi Dragotti |
MMSP | 2 |
| 2013 | Plenoptic Layer-Based Modeling for Image Based RenderingabstractImage based rendering is an attractive alternative to model based rendering for generating novel views because of its lower complexity and potential for photo-realistic results. To reduce the number of images necessary for alias-free rendering, some geometric information for the 3D scene is normally necessary. In this paper, we present a fast automatic layer-based method for synthesizing an arbitrary new view of a scene from a set of existing views. Our algorithm takes advantage of the knowledge of the typical structure of multiview data to perform occlusion-aware layer extraction. In addition, the number of depth layers used to approximate the geometry of the scene is chosen based on plenoptic sampling theory with the layers placed non-uniformly to account for the scene distribution. The rendering is achieved using a probabilistic interpolation approach and by extracting the depth layer information on a small number of key images. Numerical results demonstrate that the algorithm is fast and yet is only 0.25 dB away from the ideal performance achieved with the ground-truth knowledge of the 3D geometry of the scene of interest. This indicates that there are measurable benefits from following the predictions of plenoptic theory and that they remain true when translated into a practical system for real world data. James Pearson, Mike Brookes, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 3 |
| 2012 | Spike sorting at sub-Nyquist ratesabstractSpike sorting relies on the ability to establish the temporal occurrence of action potentials and their relation to specific neurons. Neural information is intrinsically compressible and as such suitable for sparse sampling. Potentially, this should allow for the use of multi-channel recordings, which is particularly advantageous to improve spike sorting. In this paper we propose a novel algorithm capable of sampling neural data at sub-Nyquist rates, yielding the same performance for spike sorting as traditional schemes. Jose Caballero, Jose Antonio Uriguen, Simon R. Schultz, Pier Luigi Dragotti |
ICASSP | 4 |
| 2012 | Image based rendering with depth cameras: How many are needed?abstractImage based rendering is a technique for producing arbitrary viewpoints of a scene using multiple images instead of exact object models. The recent emergence of low-price, fast, and reliable cameras for measuring depth makes possible the augmentation of traditional color images with depth images. This combination promises to improve the rendering quality of an arbitrary viewpoint and thus have a great impact on IBR. A key issue is to understand, for any particular scene of interest, how many depth images and how many color images are necessary in order to obtain good rendering results. In this paper, using a framework akin to the plenoptic function, we perform a spectral analysis of multi-view depth images in order to determine the relationship between the number of depth and color images required. Our analysis is then validated using both synthetic and real images. Christopher Gilliam, James Pearson, Mike Brookes, Pier Luigi Dragotti |
ICASSP | 4 |
| 2012 | Line-edge extraction based on E-spline acquisition model and a fast optimization algorithmabstractWe propose a line-edge extraction algorithm that uses an E-spline data acquisition model and a fast optimization algorithm. The proposed method can retrieve line-edge parameters, including orientation, offset, and amplitude at sub-pixel accuracy almost independently of the resolution of the images. Because of the optimization approach, the proposed method is robust against model mismatch such as noise, point spread function (PSF) model mismatch, or step line-edge assumption. These properties are verified by simulations using images taken with a digital SLR camera. Akira Hirabayashi, Pier Luigi Dragotti |
ICIP | 2 |
| 2012 | Video jitter analysis for automatic bootleg detectionabstractThis paper presents a novel technique for the automatic detection of recaptured videos with applications to video forensics. The proposed technique uses scene jitter as a cue for classification: when recapturing planar surfaces approximately parallel to the imaging plane, any added motion due to jitter will result in approximately uniform high-frequency 2D motion fields. The inter-frame motion trajectories are retrieved with feature tracking techniques, while local and global feature motion are decoupled through a 2-level wavelet decomposition. A normalised cross-correlation matrix is then populated with the similarities between the high-frequency components of the tracked features' trajectories. The correlation distribution is then compared with trained models for classification. Experiments with original and recaptured standard datasets show the validity of the proposed technique. Marco Visentini Scarzanella, Pier Luigi Dragotti |
MMSP | 2 |
| 2012 | Multiview Image Coding Using Depth Layers and an Optimized Bit AllocationabstractIn this paper, we present a novel wavelet-based compression algorithm for multiview images. This method uses a layer-based representation, where the 3-D scene is approximated by a set of depth planes with their associated constant disparities. The layers are extracted from a collection of images captured at multiple viewpoints and transformed using the 3-D discrete wavelet transform (DWT). The DWT consists of the 1-D disparity compensated DWT across the viewpoints and the 2-D shape-adaptive DWT across the spatial dimensions. Finally, the wavelet coefficients are quantized and entropy coded along with the layer contours. To improve the rate-distortion performance of the entire coding method, we develop a bit allocation strategy for the distribution of the available bit budget between encoding the layer contours and the wavelet coefficients. The achieved performance of our proposed scheme outperforms the state-of-the-art codecs for several data sets of varying complexity. Andriy Gelman, Pier Luigi Dragotti, Vladan Velisavljevic |
IEEE Trans. Image Process. | 2 |
| 2011 | Accurate non-iterative depth layer extraction algorithm for image based renderingabstractImage based rendering is an attractive alternative for generating novel views compared to model based rendering due to its lower complexity and potential for photo-realistic results. We present a fast unsupervised method for synthesising arbitrary viewpoints of a scene from a set of existing views. Our novel improvements include optimising the placement of depth layers to take advantage of the composition of real world scenes and hierarchically building our simple geometric model to maximise its accuracy. James Pearson, Pier Luigi Dragotti, Mike Brookes |
ICASSP | 2 |
| 2011 | Interactive multiview image codingabstractWe propose a novel multiview compression method for multiview images. The algorithm supports random access for interactive applications and has low storage requirements. The fundamental component of the method is the layer-based representation, which partitions the dataset into redundant layers characterized by a constant depth value. We exploit the redundant property of each layer and remove the side information uncertainty using Distributed Source Coding (DSC) principles. In comparison to independent coding, our method achieves a PSNR improvement of 3dB. Furthermore, we present a rate-distortion (RD) analysis which demonstrates that the proposed algorithm can achieve a better performance in comparison to independent coding. Andriy Gelman, Pier Luigi Dragotti, Vladan Velisavljevic |
ICIP | 2 |
| 2011 | Adaptive plenoptic samplingabstractThe plenoptic function enables Image-based rendering (IBR) to be viewed in terms of sampling and reconstruction. Thus the spatial sampling rate can be determined through spectral analysis of the plenoptic function. In this paper we present a method of non-uniformly sampling a scene, with a smoothly varying surface, given a finite number of samples. This method approximates such a scene with a set of slanted planes subject to the constraint of finite number of samples. We use the recent spectral analysis of a single slanted plane to determine a piecewise constant spatial sampling rate for the scene. Finally, we show that this sampling rate results in a non-uniform sampling scheme that reconstructs the plenoptic function beyond that of uniform sampling. Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes |
ICIP | 2 |
| 2010 | Multiview image compression using a layer-based representationabstractWe propose a novel compression method for multiview images. The algorithm exploits the layer-based representation, which partitions the data set into planar layers characterized by a constant depth value. For efficient compression, the partitioned data is decorrelated using the separable three-dimensional wavelet transform across the viewpoint and spatial dimensions. The transform is modified to efficiently deal with occlusions and disparity variations for different depths. The generated transform coefficients are entropy coded. Experimental results show that our coding method is capable of outperforming the state-of-the-art algorithms, like H.264/AVC, for different data sets. Andriy Gelman, Pier Luigi Dragotti, Vladan Velisavljevic |
ICIP | 2 |
| 2010 | A closed-form expression for the bandwidth of the plenoptic function under finite field of view constraintsabstractThe plenoptic function enables Image-based rendering (IBR) to be viewed in terms of sampling and reconstruction. Thus the spatial sampling rate can be determined through spectral analysis of the plenoptic function. In this paper we examine the bandwidth of the plenoptic function when both the field of view and the scene width are finite. This analysis is carried out on two planar Lambertian scenes, a fronto-parallel plane and a slanted plane, and in both cases the texture is bandlimited. We derive an exact closed-form expression for the plenoptic spectrum of a slanted plane with sinusoidal texture. We show that in both cases the finite constraints lead to band-unlimited spectra. By determining the essential bandwidth, we derive a sampling curve that gives an adequate camera spacing for a given distance between the scene and the camera line. Christopher Gilliam, Pier Luigi Dragotti, Mike Brookes |
ICIP | 2 |
| 2010 | E-spline sampling for precise and robust line-edge extractionabstractWe propose a line-edge extraction algorithm using E-spline functions as a sampling kernel. Our method is capable of extracting line-edge parameters, including amplitude, orientation, and offset, not only at sub-pixel level but also exactly provided noiseless pixel values. Even in noisy scenario, simulation results show that the proposed method outperforms a similar one based around B-spline functions with gains in standard deviation of 1.86dB for the orientation and 9.64dB for the offset when SNR is 10dB. We also show by simulations that our method extracts line-edges more precisely than the Hough transform. Akira Hirabayashi, Pier Luigi Dragotti |
ICIP | 2 |
| 2010 | Multichannel Sampling of Signals With Finite Rate of InnovationabstractIn this letter, we present a possible extension of the theory of sampling signals with finite rate of innovation (FRI) to the case of multichannel acquisition systems. The essential issue of a multichannel system is that each channel introduces different unknown delays and gains that need to be estimated for the calibration of the channels. We pose both the synchronization stage and the signal reconstruction stage as a parametric estimation problem and demonstrate that a simultaneous exact synchronization of the channels and reconstruction of the FRI signal is possible. We also consider the case of noisy measurements and evaluate the Cramér–Rao bounds (CRB) of the proposed system. Numerical results as well as the CRB show clearly that multichannel systems are more resilient to noise than the single-channel ones. Hojjat Akhondi Asl, Pier Luigi Dragotti, Loïc Baboulaz |
IEEE Signal Process. Lett. | 2 |
| 2009 | Single and multichannel sampling of bilevel polygons using exponential splinesabstractIn this paper we present a novel approach for sampling and reconstructing any K-sided convex and bilevel polygon with the use of exponential splines [1]. It will be shown that with K+1 projections we are able to perfectly reconstruct a K-sided bilevel polygon from its samples. We will also investigate the multichannel sampling scenario, consisting of a bank of E-spline filters, each with a different delay parameter compared to the reference signal. We show how by retrieving the delay parameters, we can symmetrically sample and reconstruct a given bilevel polygon using exponential splines. Hojjat Akhondi Asl, Pier Luigi Dragotti |
ICASSP | 2 |
| 2009 | Sampling signals with finite rate of innovation in the presence of noiseabstractRecently, it has been shown that it is possible to sample non-bandlimited signals that possess a limited number of degrees of freedom and uniquely reconstruct them from a finite number of uniform samples. These signals include, amongst others, streams of Diracs. In this paper, we investigate the problem of estimating the innovation parameters of a stream of Diracs from its noisy samples taken with polynomial or exponential reproducing kernels. For the one-Dirac case, we provide exact expressions for the Cramer-Rao bounds of this estimation problem. Furthermore, we propose methods to reconstruct the location of a single Dirac that reach the optimal performance given by the unbiased CramerRao bounds down to noise levels of 5 dB. Pier Luigi Dragotti, Felix Homann |
ICASSP | 1 |
| 2009 | Quadtree structured restoration algorithms for piecewise polynomial imagesabstractIterative shrinkage of sparse and redundant representations are at the heart of many state of the art denoising and deconvolution algorithms. They assume the signal is well approximated by a few elements from an overcomplete basis of a linear space. If one instead selects the elements from a nonlinear manifold it is possible to more efficiently represent piecewise polynomial signals. This suggests that image restoration algorithms based around nonlinear transformations could provide better results for this class of signals. This paper uses iterative shrinkage ideas and a nonlinear quadtree decomposition to develop image restoration algorithms suitable for piecewise polynomial images. Adam Scholefield, Pier Luigi Dragotti |
ICASSP | 2 |
| 2009 | Image restoration using a sparse quadtree decomposition representationabstractTechniques based on sparse and redundant representations are at the heart of many state of the art denoising and deconvolution algorithms. A very sparse representation of piecewise polynomial images can be obtained by using a quadtree decomposition to adaptively select a basis. We have recently exploited this to restore images of this form, however the same model can also provide very good sparse approximations of real world images. In this paper we take advantage of this to develop both image denoising and deconvolution algorithms suitable for real world images. We present results on the cameraman image showing comparable performance with iterative soft thresholding using the undecimated wavelet transform. Adam Scholefield, Pier Luigi Dragotti |
ICIP | 2 |
| 2009 | Adaptive layer extraction for image based renderingabstractImage based rendering is a promising way to produce arbitrary views of a scene using images instead of object models. However, depth variations and occlusions cause blurring in the rendered images. The solution is to use some geometrical information in order to steer the interpolation filters according to the depth. The level of detail of this geometry is often predetermined. In this paper, we present a method for extracting depth layers in the presence of occlusions for image based rendering. Moreover, we show how the layer extraction can be made to estimate depth layers in an adaptive manner, based on the spectral analysis of the plenoptic function. The rendering system therefore automatically adapts the number of depth layers based on the scene and the spacing of the sample cameras. Jesse Berent, Pier Luigi Dragotti, Mike Brookes |
MMSP | 2 |
| 2009 | Sparse sampling of structured information and its application to compressionabstractIt has been shown recently that it is possible to sample classes of non-bandlimited signals which we call signals with Finite Rate of Innovation (FRI). Perfect reconstruction is possible based on a set of suitable measurements and this provides a sharp result on the sampling and reconstruction of sparse continuous-time signals. In this paper, we first review the basic theory and results on sampling signals with finite rate of innovation. We then discuss variations of the above framework to handle noise and model mismatch. Finally, we present some results on compression of piecewise smooth signals based on the FRI framework. Pier Luigi Dragotti |
MMSP | 1 |
| 2009 | Exact Feature Extraction Using Finite Rate of Innovation Principles With an Application to Image Super-ResolutionabstractThe accurate registration of multiview images is of central importance in many advanced image processing applications. Image super-resolution, for example, is a typical application where the quality of the super-resolved image is degrading as registration errors increase. Popular registration methods are often based on features extracted from the acquired images. The accuracy of the registration is in this case directly related to the number of extracted features and to the precision at which the features are located: images are best registered when many features are found with a good precision. However, in low-resolution images, only a few features can be extracted and often with a poor precision. By taking a sampling perspective, we propose in this paper new methods for extracting features in low-resolution images in order to develop efficient registration techniques. We consider, in particular, the sampling theory of signals with finite rate of innovation and show that some features of interest for registration can be retrieved perfectly in this framework, thus allowing an exact registration. We also demonstrate through simulations that the sampling model which enables the use of finite rate of innovation principles is well suited for modeling the acquisition of images by a camera. Simulations of image registration and image super-resolution of artificially sampled images are first presented, analyzed and compared to traditional techniques. We finally present favorable experimental results of super-resolution of real images acquired by a digital camera available on the market. Loïc Baboulaz, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2009 | Geometry-Driven Distributed Compression of the Plenoptic Function: Performance Bounds and Constructive AlgorithmsabstractIn this paper, we study the sampling and the distributed compression of the data acquired by a camera sensor network. The effective design of these sampling and compression schemes requires, however, the understanding of the structure of the acquired data. To this end, we show that the a priori knowledge of the configuration of the camera sensor network can lead to an effective estimation of such structure and to the design of effective distributed compression algorithms. For idealized scenarios, we derive the fundamental performance bounds of a camera sensor network and clarify the connection between sampling and distributed compression. We then present a distributed compression algorithm that takes advantage of the structure of the data and that outperforms independent compression algorithms on real multiview images. Nicolas Gehrig, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 2 |
| 2008 | Subspace-based methods for image registration and super-resolutionabstractSuper-resolution algorithms combine multiple low resolution images into a single high resolution image. They have received a lot of attention recently in various application domains such as HDTV, satellite imaging, and video surveillance. These techniques take advantage of the aliasing present in the input images to reconstruct high frequency information of the resulting image. One of the major challenges in such algorithms is a good alignment of the input images: subpixel precision is required to enable accurate reconstruction. In this paper, we give an overview of some subspace techniques that address this problem. We first formulate super-resolution in a multichannel sampling framework with unknown offsets. Then, we present three registration methods: one approach using ideas from variable projections, one using a Fourier description of the aliased signals, and one using a spline description of the sampling kernel. The performance of the algorithms is evaluated in numerical simulations. Patrick Vandewalle, Loïc Baboulaz, Pier Luigi Dragotti, Martin Vetterli |
ICIP | 3 |
| 2007 | Unsupervised Extraction of Coherent Regions for Image Based RenderingabstractImage based rendering using undersampled light fields suffe rs from aliasing effects. These effects can be drastically reduced by usi ng some geometric information. In pop-up light field rendering [18], the scene is segmented into coherent layers, usually corresponding to approximately planar regions, that can be rendered free of aliasing. As opposed to the supervised method in the pop-up light field, we propose an unsupervised extractio n of coherent regions. The problem is posed in a multidimensional variational framework using the level set method [16]. Since the segmentation is done jointly over all the images, coherence can be imposed throughout the data. However, instead of using active hypersurfaces, we derive a semi-parametric methodology that takes into account the constraints imposed by the camera setup and the occlusion ordering. The resulting framework is a global multidimensional region competition that is consistent in all the imag es and efficiently handles occlusions. We show the validity of the method with some captured multi-view datasets. Other special effects by coherent reg ion manipulation are also demonstrated. Jesse Berent, Pier Luigi Dragotti |
BMVC | 2 |
| 2007 | Distributed Coding of Shifts using the DFT PhaseabstractIn this paper we consider the problem of image encoding with side information at the decoder, where the side information is an integer shifted version of the image at the encoder. The encoder is asked to send the shift of its own image with respect to the side information which is only available at the decoder. We propose a solution based on the encoding of the phase sign of the DFT coefficients, taken at exponentially spaced positions. We first introduce the method under ideal hypothesis, i.e. noiseless conditions without border effects, giving a theoretical foundation to the technique. Then, we consider the more realistic case of noisy images with border effects, showing the effectiveness of the proposed method. Marco Dalai, Riccardo Leonardi, Pier Luigi Dragotti |
ICASSP (1) | 3 |
| 2007 | Local Feature Extraction for Image Super-ResolutionabstractThe problem of image super-resolution from a set of low resolution multiview images has recently received much attention and can be decomposed, at least conceptually, into two consecutive steps as: registration and restoration. The ability to accurately register the input images is key to the success and the quality of image super-resolution algorithms. Using recent results from the sampling theory for signals with finite rate of innovation (FRI), we propose in this paper a new technique for subpixel extraction from low resolution images of local features like step edges and corners for image registration. By exploiting the knowledge of the sampling kernel, we are able to locate exactly the step edges on synthetic images. We also present results of full frame super-resolution of real low resolution images using our registration technique. We obtain super-resolved images with a much improved visual quality compared to using a standard local feature detection approach like a subpixel Harris corner detector. Loïc Baboulaz, Pier Luigi Dragotti |
ICIP (5) | 2 |
| 2007 | Distributed Compression of Multi-View Images using a Geometrical Coding ApproachabstractIn this paper, we propose a distributed compression approach for multi-view images, where each camera efficiently encodes its visual information locally without requiring any collaboration with the other cameras. Such a compression scheme can be necessary for camera sensor networks, where each camera has limited power and communication resources and can only transmit data to a central base station. The correlation in the multi-view data acquired by a dense multi-camera system can be extremely large and should therefore be exploited at each encoder in order to reduce the amount of data transmitted to the receiver. Our distributed source coding approach is based on a quadtree decomposition method and uses some geometrical information about the scene and the position of the cameras to estimate this multi-view correlation. We assume that the different views can be modelled as 2D piecewise polynomial functions with ID linear boundaries and show how our approach applies in this context. Our simulation results show that our approach outperforms independent encoding of real multi-view images. Nicolas Gehrig, Pier Luigi Dragotti |
ICIP (6) | 2 |
| 2006 | Distributed Sampling and Compression of Scenes with Finite Rate of Innovation in Camera Sensor NetworksabstractWe study the problem of distributed sampling and compression in sensor networks when the sensors are digital cameras that acquire a 3-D visual scene of interest from different viewing positions. We assume that sensors cannot communicate among themselves, but can process their acquired data and transmit it to a common central receiver. The main task of the receiver is then to reconstruct the best possible estimation of the original scene and the natural issue, in this context, is to understand the interplay in the reconstruction between sampling and distributed compression. In this paper, we show that if the observed scene belongs to the class of signals that can be represented with a finite number of parameters, we can determine the minimum number of sensors that allows perfect reconstruction of the scene. Then, we present a practical distributed coding approach that leads to a rate-distortion behaviour at the decoder that is independent of the number of sensors, when this number increases beyond the critical sampling. In other words, we show that the distortion at the decoder does not depend on the number of sensors used, but only on the total number of bits that can be transmitted from the sensors to the receiver Nicolas Gehrig, Pier Luigi Dragotti |
DCC | 2 |
| 2006 | Perfect Reconstruction Schemes for Sampling Piecewise Sinusoidal SignalsabstractConsider sampling a signal that is piecewise sinusoidal. Classical sampling theory does not enable a perfect reconstruction of the continuous time signal since the band is not limited (C.E. Shannon, 1949). However, we show that it is still possible to recover all the parameters of the sinusoids and the exact locations of the discontinuities using the annihilating filter method and recently developed Finite Rate of Innovation (FRI) sampling schemes (M. Vetterli et al., 2002) (P.L. Dragotti et al., 2005). Moreover, we show that there is a tradeoff between the number of sinusoids per piece and the proximity of the discontinuities in order to have a unique solution. This result recalls a sort of uncertainty principle Jesse Berent, Pier Luigi Dragotti |
ICASSP (3) | 2 |
| 2006 | Sensing and Communication With and Without BitsabstractThe successful design of sensor network architectures depends crucially on the structure of the sampling, observation, and communication processes. One of the most fundamental questions concerns the sufficiency of discrete approximations in time, space, and amplitude. In the case of space and time, the question can be rephrased as whether there is a spatio-temporal sampling theorem for typical data sets in sensor networks. This question has a positive answer in many cases of interest. The issue of discretization of amplitudes is more subtle and can be expressed as the question of whether there is a (source/channel) separation theorem for typical sensor networks. We show that this question has a negative answer in general and that the price of separation can be large. To illustrate these issues, we review the underlying theory and discuss specific examples Michael Gastpar, Martin Vetterli, Pier Luigi Dragotti |
ICASSP (5) | 3 |
| 2006 | Distributed Acquisition and Image Super-Resolution Based on Continuous Moments from SamplesabstractRecently, new sampling schemes were presented for signals with finite rate of innovation (FRI) using sampling kernels reproducing polynomials or exponentials. In this paper, we extend those sampling schemes to a distributed acquisition architecture in which numerous and randomly located sensors are pointing to the same area of interest. We emphasize the importance played by moments and show how to acquire efficiently FRI signals with a set of sensors. More importantly, we also show that those sampling schemes can be used for accurate registration of affine transformed and low-resolution images. Based on this, a new super-resolution algorithm was developed and showed good preliminary results. Loïc Baboulaz, Pier Luigi Dragotti |
ICIP | 2 |
| 2006 | Exact Local Reconstruction Algorithms for Signals with Finite Rate of InnovationabstractConsider the problem of sampling signals which are not bandlimited, but still have a finite number of degrees of freedom per unit of time, such as, for example, piecewise polynomial or piecewise sinusoidal signals, and call the number of degrees of freedom per unit of time the rate of innovation. Classical sampling theory does not enable a perfect reconstruction of such signals since they are not bandlimited. In this paper, we show that many signals with finite rate of innovation can be sampled and perfectly reconstructed using kernels of compact support and a local reconstruction algorithm. The class of kernels that we can use is very rich and includes functions satisfying strang-fix conditions, exponential splines and functions with rational Fourier transforms. Extension of such results to the 2-dimensional case are also discussed and an application to image super-resolution is presented. Pier Luigi Dragotti, Martin Vetterli, Thierry Blu |
ICIP | 1 |
| 2006 | Tomographic Approach for Sampling Multidimensional Signals with Finite Rate of InnovationabstractRecently, it was shown that it is possible to sample classes of 1-D and 2-D signal with finite rate of innovation (FRI). In particular, in P. Shukla et al., (2005), we presented local and global schemes for sampling sets of Diracs and bilevel polygons using compactly supported kernels that reproduce polynomials. In sequel to P. Shukla et al., (2005), in this paper, we present a Radon transform based hybrid scheme for sampling more general 2-D FRI signals such as piecewise polynomials with polygonal boundaries, and higher dimensional Diracs and bilevel polytopes. The key feature of the proposed scheme is an annihilating-filter-based-back-projection (AFBP) algorithm. Pancham Shukla, Pier Luigi Dragotti |
ICIP | 2 |
| 2006 | Low-Rate Reduced Complexity Image Compression using DirectionletsabstractThe standard separable two-dimensional (2-D) wavelet transform (WT) has recently achieved a great success in image processing because it provides a sparse representation of smooth images. However, it fails to capture efficiently one-dimensional (1-D) discontinuities, like edges and contours, that are anisotropic and characterized by geometrical regularity along different directions. In our previous work, we proposed a construction of critically sampled perfect reconstruction anisotropic transform with directional vanishing moments (DVM) imposed in the corresponding basis functions, called directionlets. Here, we show that the computational complexity of our transform is comparable to the complexity of the standard 2-D WT and substantially lower than the complexity of other similar approaches. We also present a zerotree-based image compression algorithm using directionlets that strongly outperforms the corresponding method based on the standard wavelets at low bit rates. Vladan Velisavljevic, Baltasar Beferull-Lozano, Martin Vetterli, Pier Luigi Dragotti |
ICIP | 4 |
| 2006 | Segmentation of Epipolar-Plane Image Volumes with Occlusion and Disocclusion CompetitionabstractConsider a dense array of cameras uniformly distributed along a line. A solid block of 3D data can be constructed by arranging the images into a stack. This volume, also known as the epipolar-plane image volume, contains highly structured data that can be segmented for object removal, insertion and compression. In this paper, we propose a segmentation scheme that takes fully advantage of the known geometry in order to model occlusions explicitly as a result of disparity. Moreover, we include this knowledge into an energy minimization scheme based on region competition with active contours. Instead of extracting layers sequentially from front to back, each layer is made to compete with the regions it is going to occlude and the ones it is going to disocclude. This enables a virtually unsupervised segmentation Jesse Berent, Pier Luigi Dragotti |
MMSP | 2 |
| 2006 | Directionlets: Anisotropic Multidirectional Representation With Separable FilteringabstractIn spite of the success of the standard wavelet transform (WT) in image processing in recent years, the efficiency of its representation is limited by the spatial isotropy of its basis functions built in the horizontal and vertical directions. One-dimensional (1-D) discontinuities in images (edges and contours) that are very important elements in visual perception, intersect too many wavelet basis functions and lead to a nonsparse representation. To efficiently capture these anisotropic geometrical structures characterized by many more than the horizontal and vertical directions, a more complex multidirectional (M-DIR) and anisotropic transform is required. We present a new lattice-based perfect reconstruction and critically sampled anisotropic M-DIR WT. The transform retains the separable filtering and subsampling and the simplicity of computations and filter design from the standard two-dimensional WT, unlike in the case of some other directional transform constructions (e.g., curvelets, contourlets, or edgelets). The corresponding anisotropic basis unctions (directionlets) have directional vanishing moments along any two directions with rational slopes. Furthermore, we show that this novel transform provides an efficient tool for nonlinear approximation of images, achieving the approximation power O(N(-1.55)), which, while slower than the optimal rate O(N(-2)), is much better than O(N(-1)) achieved with wavelets, but at similar complexity. Vladan Velisavljevic, Baltasar Beferull-Lozano, Martin Vetterli, Pier Luigi Dragotti |
IEEE Trans. Image Process. | 4 |
| 2006 | The Distributed Karhunen-Loève TransformabstractThe Karhunen-Loeve transform (KLT) is a key element of many signal processing and communication tasks. Many recent applications involve distributed signal processing, where it is not generally possible to apply the KLT to the entire signal; rather, the KLT must be approximated in a distributed fashion. This paper investigates such distributed approaches to the KLT, where several distributed terminals observe disjoint subsets of a random vector. We introduce several versions of the distributed KLT. First, a local KLT is introduced, which is the optimal solution for a given terminal, assuming all else is fixed. This local KLT is different and in general improves upon the marginal KLT which simply ignores other terminals. Both optimal approximation and compression using this local KLT are derived. Two important special cases are studied in detail, namely, the partial observation KLT which has access to a subset of variables, but aims at reconstructing them all, and the conditional KLT which has access to side information at the decoder. We focus on the jointly Gaussian case, with known correlation structure, and on approximation and compression problems. Then, the distributed KLT is addressed by considering local KLTs in turn at the various terminals, leading to an iterative algorithm which is locally convergent, sometimes reaching a global optimum, depending on the overall correlation structure. For compression, it is shown that the classical distributed source coding techniques admit a natural transform coding interpretation, the transform being the distributed KLT. Examples throughout illustrate the performance of the proposed distributed KLT. This distributed transform has potential applications in sensor networks, distributed image databases, hyper-spectral imagery, and data fusion Michael Gastpar, Pier Luigi Dragotti, Martin Vetterli |
IEEE Trans. Inf. Theory | 2 |
| 2005 | Exact sampling results for signals with finite rate of innovation using Strang-Fix conditions and local kernelsabstractRecently, it was shown that it is possible to sample classes of signals with finite rate of innovation. These sampling schemes, however, use kernels with infinite support and this leads to complex and unstable reconstruction algorithms. In this paper, we show that many signals with finite rate of innovation can be sampled and perfectly reconstructed using kernels of compact support and a local reconstruction algorithm. The class of kernels that we can use is very rich and includes any function satisfying Strang-Fix conditions, exponential splines and functions with rational Fourier transforms. Pier Luigi Dragotti, Martin Vetterli, Thierry Blu |
ICASSP (4) | 1 |
| 2005 | Dfferent - distributed and fully flexible image encoders for camera sensor networksabstractIn this paper, we propose a practical coding approach for the problem of distributed compression of multi-view images. Our coding technique is based on a tree structured compression algorithm that guarantees an optimal rate-distortion behaviour for piecewise polynomial signals. We model the different views using a piecewise polynomial function whose singularity positions are shifted from one view to the others according to the constraints imposed by the structure of the plenoptic function. We show that, starting from the optimal tree decompositions of the different views, only partial information from each tree is necessary at the decoder in order to reconstruct all the different approximations. We first present our approach in the more intuitive 1D case and show that it can be used with arbitrary bit-rate allocation. Then, we propose a construction for the case of TV different views that satisfies a certain bit conservation principle. Finally, we show how our approach can be extended to the 2D case. Nicolas Gehrig, Pier Luigi Dragotti |
ICIP (2) | 2 |
| 2005 | Sampling schemes for 2-D signals with finite rate of innovation using kernels that reproduce polynomialsabstractIn this paper, we propose new sampling schemes for classes of 2-D signals with finite rate of innovation (FRI). In particular, we consider sets of 2-D Diracs and bilevel polygons. As opposed to using only sine or Gaussian kernels [I. Maravic et al, 2004], we allow the sampling kernel to be any function that reproduces polynomials. In the proposed sampling schemes, we exploit the polynomial approximation properties of the sampling kernels in association with other relevant techniques such as complex-moments [P. Milanfar et al, 1995], annihilating filter method [M. Vetterli et al, 2002], and directional derivatives. Specifically, for the bilevel polygons, we propose two different methods: the first uses a global reconstruction algorithm and complex moments, while the second is based on directional derivatives and local reconstruction algorithms. The trade-off between these two reconstruction modalities is also briefly discussed. Pancham Shukla, Pier Luigi Dragotti |
ICIP (2) | 2 |
| 2005 | Approximation power of directionletsabstractIn spite of the success of the standard wavelet transform (WT) in image processing, the efficiency of its representation is limited by the spatial isotropy of its basis functions built in only horizontal and vertical directions. One-dimensional (1-D) discontinuities in images (edges and contours), which are very important elements in visual perception, intersect too many wavelet basis functions and reduce the sparsity of the representation. To capture efficiently these anisotropic geometrical structures, a more complex multi-directional (M-DIR) and anisotropic transform is required. We present a new lattice-based perfect reconstruction and critically sampled anisotropic M-DIR WT (with the corresponding basis functions called directionlets) that retains the separable filtering and simple filter design from the standard two-dimensional (2-D) WT and imposes directional vanishing moments (DVM). Further-more, we show that this novel transform has non-linear approximation efficiency competitive to the other previously proposed over-sampled transform constructions. Vladan Velisavljevic, Baltasar Beferull-Lozano, Martin Vetterli, Pier Luigi Dragotti |
ICIP (1) | 4 |
| 2005 | Rate-distortion optimized tree-structured compression algorithms for piecewise polynomial imagesabstractThis paper presents novel coding algorithms based on tree-structured segmentation, which achieve the correct asymptotic rate-distortion (R-D) behavior for a simple class of signals, known as piecewise polynomials, by using an R-D based prune and join scheme. For the one-dimensional case, our scheme is based on binary-tree segmentation of the signal. This scheme approximates the signal segments using polynomial models and utilizes an R-D optimal bit allocation strategy among the different signal segments. The scheme further encodes similar neighbors jointly to achieve the correct exponentially decaying R-D behavior (D(R) - c(o)2(-c1R)), thus improving over classic wavelet schemes. We also prove that the computational complexity of the scheme is of O(N log N). We then show the extension of this scheme to the two-dimensional case using a quadtree. This quadtree-coding scheme also achieves an exponentially decaying R-D behavior, for the polygonal image model composed of a white polygon-shaped object against a uniform black background, with low computational cost of O(N log N). Again, the key is an R-D optimized prune and join strategy. Finally, we conclude with numerical results, which show that the proposed quadtree-coding scheme outperforms JPEG2000 by about 1 dB for real images, like cameraman, at low rates of around 0.15 bpp. Pier Luigi Dragotti, Minh N. Do, Martin Vetterli |
IEEE Trans. Image Process. | 2 |
| 2004 | Wavelet and footprint sampling of signals with a finite rate of innovationabstractWe consider classes of not bandlimited signals, namely streams of Diracs and piecewise polynomial signals, and show that these signals can be sampled and perfectly reconstructed using wavelets as sampling kernel. Due to the multiresolution structure of the wavelet transform, these new sampling theorems naturally lead to the development of a new resolution enhancement algorithm based on wavelet footprints (Dragotti, P.L. and Vetterli, M., IEEE Trans. Sig. Process., vol.51, no.5, p.1306-23, 2003). Preliminary results show the potentiality of this algorithm. Pier Luigi Dragotti, Martin Vetterli |
ICASSP (2) | 1 |
| 2004 | On compression using the distributed Karhunen-Loeve transformabstractIn this paper, we discuss a framework for the distributed compression of vector sources, based on our previous work on distributed transform coding (2002, 2003). In particular, our goal is to develop a strategy of first applying a suitable distributed Karhunen-Loeve transform, whereafter each component can be handled by standard distributed compression techniques. In the present paper, we first study the scenario where all but one terminal furnish a noisy approximation of their observation. For the case where the underlying vector is Gaussian, and the added noise is also Gaussian, we establish that indeed it is optimal for the last terminal to apply a (local) transform to its observations, and to separately compress each component in the transform domain. Then, we outline how this leads to a general simple distributed compression strategy for Gaussian vector sources: each terminal applies a suitable local transform to its observations, and encodes the resulting components separately in a Wyner-Ziv fashion, i.e., treating the compressed descriptions of all other terminals as side information available to the decoder. This achieves the best known performance. The optimum performance in unknown to date. Michael Gastpar, Pier Luigi Dragotti, Martin Vetterli |
ICASSP (3) | 2 |
| 2004 | Distributed compression of the plenoptic functionabstractIn this paper, we consider the problem of distributed compression in camera sensor networks. Due to the spatial proximity of the different cameras, acquired images can be highly dependent. The correlation in the visual information retrieved is related to the structure of the plenoptic function and can be estimated using geometrical information such as the position of the cameras and some bounds on the location of the objects. We propose a distributed compression scheme that takes advantage of this geometrical information in order to reduce the overall transmission rate from the sensors to a common central receiver. This new approach allows for a flexible repartition of the transmission bit-rates amongst the encoders and is optimal in many cases. Moreover, we show that our coding scheme can be made resilient to a fixed number of occlusions and that perfect reconstruction and interpolation are possible at the receiver. Nicolas Gehrig, Pier Luigi Dragotti |
ICIP | 2 |
| 2004 | Distributed compression in camera sensor networksabstractWe address the problem of distributed compression in camera sensor networks. Our approach uses some geometrical information in order to estimate the correlation in the visual data. This correlation, which is related to the structure of the plenoptic function, can then be used to reduce the overall transmission rate from the sensors to a common central receiver. Our approach allows for a flexible allocation of the bit-rates amongst the encoders and can be made resilient to a fixed number of occlusions. Finally, we show that our distributed coding approach can be extended to general binary sources. The technique we propose uses linear channel codes and can achieve any point of the Slepian-Wolf achievable rate region. Nicolas Gehrig, Pier Luigi Dragotti |
MMSP | 2 |
| 2003 | The Distributed, Partial, And Conditional Karhunen-Loève TransformsabstractThe Karhunen-Loeve transform (KLT) is a key element of many signal processing tasks, including approximation, compression, and classification. Many recent applications involve distributed signal processing where it is not generally possible to apply the KLT to the signal; the KLT must be approximated in a distributed fashion. Investigations were carried out on the distributed approximations to the KLT. First, explicit solutions to special cases were presented including a partial KLT, a conditional KLT, and the combination of these two special cases. These results were used to derive an algorithm that finds the best distributed approximation to the KLT. Applications of the results from sensor networks and distributed databases were discussed. Michael Gastpar, Pier Luigi Dragotti, Martin Vetterli |
DCC | 2 |
| 2003 | Distributed signal processing and communications: on the interaction of sources and channelsabstractDistributed ways of communicating, processing, and sensing are replacing more traditional centralized architectures. An early example of this revolution in distributed communications is appearing in the form of sensor networks, which are densely distributed networks of embedded signal sensors, controls and processors. These nodes could be simple signal sensors, but could also be cameras and microphones. In this distributed scenario, there are several interesting topics to investigate that span from traditional signal processing problems (i.e. sampling, compression, detection) to communication and information theory (i.e. transmission protocols, capacity bounds for ad-hoc networks). This paper reviews some recent results on the topic of source representations and distributed source coding and transmission. Thibaut Ajdler, Razvan Cristescu, Pier Luigi Dragotti, Michael Gastpar, Irena Maravic, Martin Vetterli |
ICASSP (4) | 3 |
| 2003 | Sampling and interpolation of the plenoptic functionabstractIn this paper, we present reconstruction schemes for the plenoptic function. Using new sampling methods, we show that, for some particular scenes, it is possible to perfectly reconstruct the plenoptic function from a finite number of cameras with finite resolution (Theorem 1 and Corollary 1). In more general cases, we demonstrate new ways of interpolating exactly the plenoptic function (Theorem 2). Finally, we show that it is possible to infer camera locations from a finite set of images. In all cases, we have perfect solutions due to the super-resolution property of the sampling. First numerical experiments on noisy observations show the potentiality of this new theoretical developments. Amina Chebira, Pier Luigi Dragotti, Luciano Sbaiz, Martin Vetterli |
ICIP (2) | 2 |
| 2003 | Filter banks for multiple description codingabstractIn this paper we review some of our recent results on the design of critically sampled and oversampled filter banks for multiple description coding. For the case of critically sampled filter banks, we show that optimal filters are obtained by allocating the redundancy over frequency with a reverse 'water-filling' strategy. Then we present families of oversampled filter banks that attain optimal performance as well. Pier Luigi Dragotti |
ICIP (1) | 1 |
| 2003 | Discrete multidirectional wavelet basesabstractThe application of the wavelet transform in image processing is most frequently based on a separable construction. While simple, such an approach is not capable of capturing properly all 2D properties in images. In this paper, a new truly separable multidirectional transform is proposed with a subsampling method based on lattice theory. Applications are possible in many areas of image processing. Some promising improvements are achieved in nonlinear approximation and denoising of images. Vladan Velisavljevic, Baltasar Beferull-Lozano, Martin Vetterli, Pier Luigi Dragotti |
ICIP (1) | 4 |
| 2003 | Discrete directional wavelet bases for image compression
Pier Luigi Dragotti, Vladan Velisavljevic, Martin Vetterli, Baltasar Beferull-Lozano |
VCIP | 1 |
| 2002 | Deconvolution with wavelet footprints for ill-posed inverse problemsabstractIn recent years, wavelet based algorithms have been successful in different signal processing tasks. The wavelet transform is a powerful tool, because it manages to efficiently represent sharp discontinuities. Indeed, discontinuities carry most of the signal information and, so, they represent the most critical part to analyse. We have recently introduced the notion of footprints, which form an overcomplete basis built on the wavelet transform. With footprints, one can exactly model the dependency across scales of the wavelet coefficients generated by a discontinuity and this allows to further improve wavelet based algorithms. In this paper we present a footprint based algorithm for signal deconvolution. The algorithm is fast and works for blind deconvolution too. With footprints we manage to deconvolve efficiently the irregular part of the signal. Thanks to the property of footprints of exactly modeling discontinuities, the deconvolved signal does not present artifacts around discontinuities. Moreover, we show that the residual, that is, the difference between the deconvolved signal with footprints and the observed signal, is regular. Thus, this residual can be further deconvolved with any other traditional method. We show that our system outperforms other deconvolution methods. Pier Luigi Dragotti, Martin Vetterli |
ICASSP | 1 |
| 2002 | Directional wavelet transforms and framesabstractThe application of the wavelet transform in image processing is most frequently based on a separable transform. Lines and columns in an image are treated independently and the basis functions are simply products of corresponding one-dimensional functions. Such a method keeps simplicity in design and computation. A new two-dimensional approach is proposed, which retains the simplicity of separable processing, but allows more directionalities. The method can be applied in many areas like denoising, nonlinear approximation and compression. The results on nonlinear approximation and denoising show interesting gains compared to the standard two-dimensional analysis. Pier Luigi Dragotti, Vladan Velisavljevic, Martin Vetterli |
ICIP (3) | 1 |
| 2002 | Improved quadtree algorithm based on joint coding for piecewise smooth image compressionabstractWe present a novel coding algorithm based on the tree structured segmentation, which achieves oracle like rate-distortion (R-D) behavior for a simple class of signals, namely piecewise polynomials in the high bit rate regime. We consider a R-D optimization framework, which employs optimal bit allocation strategy among different signal segments to achieve the best tradeoff between description complexity and approximation quality. First, we describe the basic idea of the algorithm for the 1D case. It can be shown that the proposed compression algorithm based on an optimal binary tree segmentation achieves the oracle like R-D behavior (D(R)/spl sim/c/sub 0/2/sup -c1R/) with the computational cost of the order O(NlogN). We then show the extension of the scheme to the 2D case with the similar R-D behavior without sacrificing the computational ease. Finally, we conclude with some experimental results. Pier Luigi Dragotti, Minh N. Do, Martin Vetterli |
ICME (1) | 2 |
| 2002 | Rate-distortion optimized tree based coding algorithmsabstractThis paper addresses the problem of efficient coding of an important class of signals, namely piecewise polynomials. For this signal class, we develop a coding algorithm, which achieves oracle like rate-distortion (R-D) behavior in the high bit rate regime and with a reasonable computational complexity. For the 1-D case, our scheme is based on the binary tree segmentation of the signal and an optimal bit allocation strategy among the different signal segments. The scheme further encodes the similar neighbors jointly to achieve the right exponentially decaying R-D behavior (D(R) /spl sim/ c/sub 0/2/sup -c1R/). We have also shown that the computational cost of the scheme is of the order O(N log N). We then show that the scheme can be easily extended to the 2-D case, as the quadtree based coding scheme, with the similar R-D behavior and computational cost. Finally, we conclude with some numerical results. Pier Luigi Dragotti, Minh N. Do, Martin Vetterli |
ITW | 2 |
| 2002 | Optimal filter banks for multiple description coding: Analysis and synthesisabstractMultiple description (MD) coding is a source coding technique for information transmission over unreliable networks. In MD coding, the coder generates several different descriptions of the same signal and the decoder can produce a useful reconstruction of the source with any received subset of these descriptions. In this paper, we study the problem of MD coding of stationary Gaussian sources with memory. First, we compute an approximate MD rate distortion region for these sources, which we prove to be asymptotically tight at high rates. This region generalizes the MD rate distortion region of El Gamal and Cover (1982), and Ozarow (1980) for memoryless Gaussian sources. Then, we develop an algorithm for the design of optimal two-channel biorthogonal filter banks for MD coding of Gaussian sources. We show that optimal filters are obtained by allocating the redundancy over frequency with a reverse "water-filling" strategy. Finally, we present experimental results which show the effectiveness of our filter banks in the low complexity, low rate regime. Pier Luigi Dragotti, Sergio D. Servetto, Martin Vetterli |
IEEE Trans. Inf. Theory | 1 |
| 2002 | Filter bank frame expansions with erasuresabstractWe study frames for robust transmission over the Internet. In our previous work, we used quantized finite-dimensional frames to achieve resilience to packet losses; here, we allow the input to be a sequence in l/sub 2/(Z) and focus on a filter-bank implementation of the system. We present results in parallel, R/sup N/ or C/sup N/ versus l/sub 2/(Z), and show that uniform tight frames, as well as newly introduced strongly uniform tight frames, provide the best performance. Jelena Kovacevic, Pier Luigi Dragotti, Vivek K. Goyal |
IEEE Trans. Inf. Theory | 2 |
| 2001 | Quantized Oversampled Filter Banks with ErasuresabstractOversampled filter banks can be used to enhance resilience to erasures in communication systems in much the same way that finite-dimensional frames have previously been applied. This paper extends previous finite dimensional treatments to frames and signals in l/sub 2/(Z) with frame expansions that can be implemented efficiently with filter banks. It is shown that tight frames attain best performance. In particular, if encoding with a uniform frame, the quantization error is minimized if and only if the frame is tight. In case of one erasure and if encoding with a strongly uniform frame, tight frames are still optimal. In case of more erasures, an expression for the mean square error is given and some general considerations are presented. Pier Luigi Dragotti, Jelena Kovacevic, Vivek K. Goyal |
Data Compression Conference | 1 |
| 2001 | On the compression of two-dimensional piecewise smooth functionsabstractIt is well known that wavelets provide good non-linear approximation of one-dimensional (1-D) piecewise smooth functions. However, it has been shown that the use of a basis with good approximation properties does not necessarily lead to a good compression algorithm. The situation in 2-D is much more complicated since wavelets are not good for modeling piecewise smooth signals (where discontinuities are along smooth curves). The purpose of this work is to analyze the performance of compression algorithms for 2-D piecewise smooth functions directly in a rate distortion context. We consider some simple image models and compute rate distortion bounds achievable using oracle based methods. We then present a practical compression algorithm based on optimal quadtree decomposition that, in some cases, achieve the oracle performance. Pier Luigi Dragotti, Minh N. Do, Martin Vetterli |
ICIP (1) | 1 |
| 2001 | Footprints and edgeprints for image denoising and compressionabstractWavelets have been quite successful in compression or denoising applications. To further improve the performance of wavelet based algorithms, we have recently introduced the notion of footprint, which is a data structure which contains all the wavelet coefficients generated by a discontinuity. The combined use of wavelets and footprints leads to very efficient algorithms for compression and denoising of 1D piecewise smooth signals. We extend some of the previous results by presenting a new denoising algorithm, where footprints are chosen adaptively according to the singularity locations. This new algorithm outperforms previously proposed ones. Then, we introduce the notion of edgeprints, which represents a natural extension of footprints to the two dimensional case. First experimental results on the compression of 2D piecewise smooth signals using edgeprints are promising. Pier Luigi Dragotti, Martin Vetterli |
ICIP (2) | 1 |
| 2000 | Analysis of Optimal Filter Banks for Multiple Description CodingabstractWe study the problem of multiple description (MD) coding of stationary Gaussian sources with memory. First, we compute an approximate rate distortion region for these sources, which we prove to be asymptotically tight at high rates: this region generalizes the standard MD rate distortion region for memoryless sources. Then we develop an algorithm for the design of optimal biorthogonal filter banks for MD coding. Finally. We present some experimental results, where we measure the deviation from optimality of our proposed system. For almost uncorrelated sources the gap between the performance of our proposed system and the ideal bounds is quite high, on the other hand for highly correlated sources this gap is reduced, due to the ability of our system to take advantage of the memory in the source. In this case, in realistic scenarios where finite complexity/delay is an issue, the subband coding approach is competitive with other approaches such as the decorrelating transform followed by MD scalar quantizers. Pier Luigi Dragotti, Sergio D. Servetto, Martin Vetterli |
Data Compression Conference | 1 |
| 2000 | Wavelet Transform Footprints: Catching Singularities for Compression and DenoisingabstractWavelets have been widely used for signal compression, image compression being a prime example, and for signal denoising. What makes wavelets such an attractive tool is their capability of representing both transient and stationary behaviors of a signal with few coefficients. We consider the problem of compressing and denoising a particular class of functions: piecewise polynomial signals. We show the limit of usual wavelet coders and present an alternative compression algorithm. The main innovation of the algorithm is that it tries to efficiently compress the significant coefficients of the wavelet decomposition rather then the zero coefficients as in usual coders. The proposed algorithm can potentially be extended to more general signals and represents an effective solution to problems like signal denoising and image compression. Pier Luigi Dragotti, Martin Vetterli |
ICIP | 1 |
| 2000 | Compression of multispectral images by three-dimensional SPIHT algorithmabstractThe authors carry out low bit-rate compression of multispectral images by means of the Said and Pearlman's SPIHT algorithm, suitably modified to take into account the interband dependencies. Two techniques are proposed: in the first, a three-dimensional (3D) transform is taken (wavelet in the spatial domain, Karhunen-Loeve in the spectral domain) and a simple 3D SPIHT is used; in the second, after taking a spatial wavelet transform, spectral vectors of pixels are vector quantized and a gain-driven SPIHT is used. Numerous experiments on two sample multispectral images show very good performance for both algorithms. Pier Luigi Dragotti, Giovanni Poggi, Arturo R. P. Ragozini |
IEEE Trans. Geosci. Remote. Sens. | 1 |