Nikos Deligiannis

dblp:90/5258 · DBLP profile ↗
← Back
90ranked-venue papers
12as first author
36since 2021 · last 2026
0000-0001-9300-5860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Computer networks · 10 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Power-Efficient Spiking Conversion of Deep Unfolded Transformers
Ahmed Sadaqa, Brent De Weerdt, Ruiqi Chen 0001, Nikos Deligiannis, Bruno da Silva 0001
ISCAS4
2025 Zero-Shot Natural Language Explanations
abstract
Natural Language Explanations (NLEs) interpret the decision-making process of a given model through textual sentences. Current NLEs suffer from a severe limitation; they are unfaithful to the model’s actual reasoning process, as a separate textual decoder is explicitly trained to generate those explanations using annotated datasets for a specific task, leading them to reflect what annotators desire. In this work, we take the first step towards generating faithful NLEs for any visual classification model without any training data. Our approach models the relationship between class embeddings from the classifier of the vision model and their corresponding class names via a simple MLP which trains in seconds. After training, we can map any new text to the classifier space and measure its association with the visual features. We conduct experiments on 38 vision models, including both CNNs and Transformers. In addition to NLEs, our method offers other advantages such as zero-shot image classification and fine-grained concept discovery.
Fawaz Sammani, Nikos Deligiannis
ICLR2
2024 Unveiling Privacy Risks in Stochastic Neural Networks Training: Effective Image Reconstruction from Gradients
Nikos Deligiannis
ECCV (30)3
2024 Physics-Guided Variational Graph Autoencoder For Air Quality Inference
abstract
Urban air quality modelling aims at inferring unknown pollution concentrations at specific urban locations. Physical methods derive the partial differential equations (PDEs) that mathematically define the laws of motion, albeit using computationally intense algorithms. In contrast, deep (generative) models, such as variational autoencoders, provide high performance by addressing the task as a data generation problem. Yet, physics knowledge and the spatio-temporal data correlations are not exploited by these deep learning models. In this work, we propose a physics-guided variational graph autoencoder whose graph convolutional operator is derived from the PDE defining the convection-diffusion physical process. We compare against statistical and deep learning approaches on two air quality datasets and report superior performance.
Esther Rodrigo Bonet, Nikos Deligiannis
ICASSP2
2024 Learned Layered Coding for Successive Refinement in the Wyner-Ziv Problem
abstract
We propose a data-driven approach to explicitly learn the progressive encoding of a continuous source, which is successively decoded with increasing levels of quality and with the aid of correlated side information. This setup refers to the successive refinement of the Wyner-Ziv coding problem. Assuming ideal Slepian-Wolf coding, our approach employs recurrent neural networks (RNNs) to learn layered encoders and decoders for the quadratic Gaussian case. The models are trained by minimizing a variational bound on the rate-distortion function of the successively refined Wyner-Ziv coding problem. We demonstrate that RNNs can explicitly retrieve layered binning solutions akin to scalable nested quantization. Moreover, the rate-distortion performance of the scheme is on par with the corresponding monolithic Wyner-Ziv coding approach and is close to the rate-distortion bound.
Boris Joukovsky, Brent De Weerdt, Nikos Deligiannis
ICASSP3
2024 Co-Occurrence Graph-Enhanced Hierarchical Prediction of ICD Codes
abstract
Recent healthcare applications of natural language processing involve multi-label classification of health records using the International Classification of Diseases (ICD). While prior research highlights intricate text models and explores external knowledge like hierarchical ICD ontology, fewer studies integrate code relationships from whole datasets to enhance ICD coding accuracy. This study presents a modular approach, sequentially combining graph-based integration of ICD code co-occurrence with a hard-coded hierarchical-enriched text representation drawn from the ICD ontology. Findings reveal: 1) significant performance gains in the combined model, aside from the significant performance gain in each enhancement module in isolation, 2) graph-based module’s efficacy is more pronounced when applied to enhanced features using the hierarchical ICD ontology, and 3) experiments demonstrate hierarchy depth’s impact on performance, concluding the deepest level’s enrichment.
Soha Sadat Mahdi, Eirini Papagiannopoulou, Nikos Deligiannis, Hichem Sahli
ICASSP3
2024 On The Detection Of Images Generated From Text
abstract
The introduction of diverse text-to-image generation models has sparked significant interest across various sectors. While these models provide the groundbreaking capability to convert textual descriptions into visual data, their widespread usage has ignited concerns over misusing realistic synthesized images. Despite the pressing need, research on detecting such synthetic images remains limited. This paper aims to bridge this gap by evaluating the ability of several existing detectors to detect synthesized images produced by text-to-image generation models. Our research includes testing four popular text-to-image generation models: Stable Diffusion (SD), Latent Diffusion (LD), GLIDE, and DALL$\cdot$E-MINI (DM), and leverages two benchmark prompt-image datasets as real images. Additionally, our research focuses on identifying robust, efficient, lightweight detectors to minimize computational resource usage. Recognizing the limitations of current detection approaches, we propose a novel detector grounded in latent space analysis tailored for recognizing text-to-image synthesized visuals. Experimental results demonstrate that the proposed detector not only achieves high prediction accuracy but also exhibits enhanced robustness against image perturbations while maintaining lower computational complexity compared to existing models in detecting text-to-image generated synthetic images.
Charuka Moremada, Nikos Deligiannis
ICIP3
2024 Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge
abstract
Contrastive Language-Image Pretraining (CLIP) performs zero-shot image classification by mapping images and textual class representation into a shared embedding space, then retrieving the class closest to the image. This work provides a new approach for interpreting CLIP models for image classification from the lens of mutual knowledge between the two modalities. Specifically, we ask: what concepts do both vision and language CLIP encoders learn in common that influence the joint embedding space, causing points to be closer or further apart? We answer this question via an approach of textual concept-based explanations, showing their effectiveness, and perform an analysis encompassing a pool of 13 CLIP models varying in architecture, size and pretraining datasets. We explore those different aspects in relation to mutual knowledge, and analyze zero-shot predictions. Our approach demonstrates an effective and human-friendly way of understanding zero-shot classification decisions with CLIP.
Fawaz Sammani, Nikos Deligiannis
NeurIPS2
2024 Holistic Representation Learning for Multitask Trajectory Anomaly Detection
abstract
Video anomaly detection deals with the recognition of abnormal events in videos. Apart from the visual signal, video anomaly detection has also been addressed with the use of skeleton sequences. We propose a holistic representation of skeleton trajectories to learn expected motions across segments at different times. Our approach uses multitask learning to reconstruct any continuous unobserved temporal segment of the trajectory allowing the extrapolation of past or future segments and the interpolation of in-between segments. We use an end-to-end attention-based encoder-decoder. We encode temporally occluded trajectories, jointly learn latent representations of the occluded segments, and reconstruct trajectories based on expected motions across different temporal segments. Extensive experiments on three trajectory-based video anomaly detection datasets show the advantages and effectiveness of our approach with state-of-the-art results on anomaly detection in skeleton trajectories†.
Alexandros Stergiou, Brent De Weerdt, Nikos Deligiannis
WACV3
2024 Semi-supervised medical image classification via distance correlation minimization and graph attention regularization
Abel Díaz Berenguer, Maryna Kvasnytsia, Matías N. Bossa, Tanmoy Mukherjee, Nikos Deligiannis, Hichem Sahli
Medical Image Anal.5
2024 The STOIC2021 COVID-19 AI challenge: Applying reusable training methodologies to private data
abstract
Challenges drive the state-of-the-art of automated medical image analysis. The quantity of public training data that they provide can limit the performance of their solutions. Public access to the training methodology for these solutions remains absent. This study implements the Type Three (T3) challenge format, which allows for training solutions on private data and guarantees reusable training methodologies. With T3, challenge organizers train a codebase provided by the participants on sequestered training data. T3 was implemented in the STOIC2021 challenge, with the goal of predicting from a computed tomography (CT) scan whether subjects had a severe COVID-19 infection, defined as intubation or death within one month. STOIC2021 consisted of a Qualification phase, where participants developed challenge solutions using 2000 publicly available CT scans, and a Final phase, where participants submitted their training methodologies with which solutions were trained on CT scans of 9724 subjects. The organizers successfully trained six of the eight Final phase submissions. The submitted codebases for training and running inference were released publicly. The winning solution obtained an area under the receiver operating characteristic curve for discerning between severe and non-severe COVID-19 of 0.815. The Final phase solutions of all finalists improved upon their Qualification phase solutions.
Luuk H. Boulogne, Julian Lorenz, Daniel Kienzle, Robin Schön, Katja Ludwig, Rainer Lienhart, Simon Jégou, Derik Shi, Mayug Maniparambil, Dominik Müller, Silvan Mertes, Niklas Schröter, Fabio Hellmann, Miriam Elia, Ine Dirks, Matías N. Bossa, Abel Díaz Berenguer, Tanmoy Mukherjee, Jef Vandemeulebroucke, Hichem Sahli, Nikos Deligiannis, Panagiotis Gonidakis, Ngoc Dung Huynh, Muhammad Imran Razzak, Mohamed Reda Bouadjenek, Mario Verdicchio, Pasquale Borrelli, Marco Aiello 0003, James A. Meakin, Alexander Lemm, Christoph Russ, Razvan Ionasec, Nikos Paragios, Bram van Ginneken, Marie-Pierre Revel
Medical Image Anal.24
2024 Interpretable Neural Networks for Video Separation: Deep Unfolding RPCA With Foreground Masking
abstract
We present two deep unfolding neural networks for the simultaneous tasks of background subtraction and foreground detection in video. Unlike conventional neural networks based on deep feature extraction, we incorporate domain-knowledge models by considering a masked variation of the robust principal component analysis problem (RPCA). With this approach, we separate video clips into low-rank and sparse components, respectively corresponding to the backgrounds and foreground masks indicating the presence of moving objects. Our models, coined ROMAN-S and ROMAN-R, map the iterations of two alternating direction of multipliers methods (ADMM) to trainable convolutional layers, and the proximal operators are mapped to non-linear activation functions with trainable thresholds. This approach leads to lightweight networks with enhanced interpretability that can be trained on limited data. In ROMAN-S, the correlation in time of successive binary masks is controlled with side-information based on$\ell _{1}$-$\ell _{1}$minimization. ROMAN-R enhances the foreground detection by learning a dictionary of atoms to represent the moving foreground in a high-dimensional feature space and by using reweighted-$\ell _{1}$-$\ell _{1}$minimization. Experiments are conducted on both synthetic and real video datasets, for which we also include an analysis of the generalization to unseen clips. Comparisons are made with existing deep unfolding RPCA neural networks, which do not use a mask formulation for the foreground, and with a 3D U-Net baseline. Results show that our proposed models outperform other deep unfolding networks, as well as the untrained optimization algorithms. ROMAN-R, in particular, is competitive with the U-Net baseline for foreground detection, with the additional advantage of providing video backgrounds and requiring substantially fewer training parameters and smaller training sets.
Boris Joukovsky, Yonina C. Eldar, Nikos Deligiannis
IEEE Trans. Image Process.3
2024 Visualizing and Understanding Contrastive Learning
abstract
Contrastive learning has revolutionized the field of computer vision, learning rich representations from unlabeled data, which generalize well to diverse vision tasks. Consequently, it has become increasingly important to explain these approaches and understand their inner workings mechanisms. Given that contrastive models are trained with interdependent and interacting inputs and aim to learn invariance through data augmentation, the existing methods for explaining single-image systems (e.g., image classification models) are inadequate as they fail to account for these factors and typically assume independent inputs. Additionally, there is a lack of evaluation metrics designed to assess pairs of explanations, and no analytical studies have been conducted to investigate the effectiveness of different techniques used to explaining contrastive learning. In this work, we design visual explanation methods that contribute towards understanding similarity learning tasks from pairs of images. We further adapt existing metrics, used to evaluate visual explanations of image classification systems, to suit pairs of explanations and evaluate our proposed methods with these metrics. Finally, we present a thorough analysis of visual explainability methods for contrastive learning, establish their correlation with downstream tasks and demonstrate the potential of our approaches to investigate their merits and drawbacks.
Fawaz Sammani, Boris Joukovsky, Nikos Deligiannis
IEEE Trans. Image Process.3
2024 Long-Term Regional Influenza-Like-Illness Forecasting Using Exogenous Data
abstract
Disease forecasting is a longstanding problem for the research community, which aims at informing and improving decisions with the best available evidence. Specifically, the interest in respiratory disease forecasting has dramatically increased since the beginning of the coronavirus pandemic, rendering the accurate prediction of influenza-like-illness (ILI) a critical task. Although methods for short-term ILI forecasting and nowcasting have achieved good accuracy, their performance worsens at long-term ILI forecasts. Machine learning models have outperformed conventional forecasting approaches enabling to utilize diverse exogenous data sources, such as social media, internet users' search query logs, and climate data. However, the most recent deep learning ILI forecasting models use only historical occurrence data achieving state-of-the-art results. Inspired by recent deep neural network architectures in time series forecasting, this work proposes the Regional Influenza-Like-Illness Forecasting (ReILIF) method for regional long-term ILI prediction. The proposed architecture takes advantage of diverse exogenous data, that are, meteorological and population data, introducing an efficient intermediate fusion mechanism to combine the different types of information with the aim to capture the variations of ILI from various views. The efficacy of the proposed approach compared to state-of-the-art ILI forecasting methods is confirmed by an extensive experimental study following standard evaluation measures.
Eirini Papagiannopoulou, Matías N. Bossa, Nikos Deligiannis, Hichem Sahli
IEEE J. Biomed. Health Informatics3
2024 SNIPPET: A Framework for Subjective Evaluation of Visual Explanations Applied to DeepFake Detection
abstract
Explainable Artificial Intelligence (XAI) attempts to help humans understand machine learning decisions better and has been identified as a critical component toward increasing the trustworthiness of complex black-box systems, such as deep neural networks. In this article, we propose a generic and comprehensive framework named SNIPPET and create a user interface for the subjective evaluation of visual explanations, focusing on finding human-friendly explanations. SNIPPET considers human-centered evaluation tasks and incorporates the collection of human annotations. These annotations can serve as valuable feedback to validate the qualitative results obtained from the subjective assessment tasks. Moreover, we consider different user background categories during the evaluation process to ensure diverse perspectives and comprehensive evaluation. We demonstrate SNIPPET on a DeepFake face dataset. Distinguishing real from fake faces is a non-trivial task even for humans that depends on rather subtle features, making it a challenging use case. Using SNIPPET, we evaluate four popular XAI methods which provide visual explanations: Gradient-weighted Class Activation Mapping, Layer-wise Relevance Propagation, attention rollout, and Transformer Attribution. Based on our experimental results, we observe preference variations among different user categories. We find that most people are more favorable to the explanations of rollout. Moreover, when it comes to XAI-assisted understanding, those who have no or lack relevant background knowledge often consider that visual explanations are insufficient to help them understand. We open-source our framework for continued data collection and annotation at https://github.com/XAI-SubjEvaluation/SNIPPET .
Boris Joukovsky, José Oramas M., Tinne Tuytelaars, Nikos Deligiannis
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Designing Transformer Networks for Sparse Recovery of Sequential Data Using Deep Unfolding
abstract
Deep unfolding models are designed by unrolling an optimization algorithm into a deep learning network. These models have shown faster convergence and higher performance compared to the original optimization algorithms. Additionally, by incorporating domain knowledge from the optimization algorithm, they need much less training data to learn efficient representations. Current deep unfolding networks for sequential sparse recovery consist of recurrent neural networks (RNNs), which leverage the similarity between consecutive signals. We redesign the optimization problem to use correlations across the whole sequence, which unfolds into a Transformer architecture. Our model is used for the task of video frame reconstruction from low-dimensional measurements and is shown to outperform state-of-the-art deep unfolding RNN and Transformer models, as well as a traditional Vision Transformer on several video datasets.
Brent De Weerdt, Yonina C. Eldar, Nikos Deligiannis
ICASSP3
2023 Relevance Propagation through Deep Conditional Random Fields
abstract
Conditional random fields (CRFs), a particular type of graph neural networks (GNNs), can be used to make structured predictions in machine learning, with various applications from image processing and natural language processing to recommender systems. CRFs refine the prediction of a sample by taking into account its context information. However, there is a lack of work on post-hoc explanation approaches to CRFs, especially when the model is softmax-activated like the deep mean field network (DMFN). In this paper, we bridge this gap by proposing a layer-wise relevance propagation (LRP) method based on deep Taylor decomposition to explain CRFs, especially the DMFN model. The method considers the intermediate softmax activation layers in DMFN. We use two evaluation settings: top K% deletion and insertion to evaluate the method. Experimental studies on fake news detection using the DMFN model prove the effectiveness of our explanation method compared to the other baseline methods.
Boris Joukovsky, Nikos Deligiannis
ICASSP3
2023 Leaping Into Memories: Space-Time Deep Feature Synthesis
abstract
The success of deep learning models has led to their adaptation and adoption by prominent video understanding methods. The majority of these approaches encode features in a joint space-time modality for which the inner workings and learned representations are difficult to visually interpret. We propose LEArned Preconscious Synthesis (LEAPS), an architecture-independent method for synthesizing videos from the internal spatiotemporal representations of models. Using a stimulus video and a target class, we prime a fixed space-time model and iteratively optimize a video initialized with random noise. Additional regularizers are used to improve the feature diversity of the synthesized videos alongside the cross-frame temporal coherence of motions. We quantitatively and qualitatively evaluate the applicability of LEAPS by inverting a range of spatiotemporal convolutional and attention-based architectures trained on Kinetics-400, which to the best of our knowledge has not been previously accomplished.1
Alexandros Stergiou, Nikos Deligiannis
ICCV2
2023 Locally Accumulated Adam For Distributed Training With Sparse Updates
abstract
The high bandwidth required for gradient exchange is a bottleneck for the distributed training of large transformer models. Most sparsification approaches focus on gradient compression for convolutional neural networks (CNNs) optimized by SGD. In this work, we show that performing local gradient accumulation when using Adam to optimize transformers in distributed fashion leads to a misled optimization direction and we address this problem by accumulating the optimization direction locally. We also empirically demonstrate most sparse gradients do not overlap and thus show that sparsification is comparable to an asynchronous update. Our experiments with classification and segmentation tasks show that our method can still maintain the correct optimization direction in distributed training event under highly sparse updates.
Nikos Deligiannis
ICIP2
2023 Model-Agnostic Visual Explanations via Approximate Bilinear Models
abstract
This paper proposes InteractionLIME: a model-agnostic attribution technique to explain deep models predictions in terms of feature interactions. Specifically, we regress a bilinear form to approximate the output of two-input models, by sampling perturbations of both inputs simultaneously. Upon training, we retrieve a global explanation and a set of feature partitioning maps via the singular value decomposition of the learned interaction matrix of the bilinear model. We demonstrate InteractionLIME on vision and text-vision contrastive models, using visual examples and quantitative evaluation metrics. Our results show that the bilinear model successfully retrieves important interacting features from both inputs, while strongly reducing the occurrence of incomplete or asymmetric explanations produced by a linear model.
Boris Joukovsky, Fawaz Sammani, Nikos Deligiannis
ICIP3
2023 KORDI: A Framework for Real-Time Performance and Cost Optimization of Apache Spark Streaming
abstract
Apache Spark is one of the most commonly used frameworks for Big Data processing. Research on the provided streaming dynamic resource allocation feature, has been shown that large data load fluctuations, for instance, in website traffic, have a negative impact on the automatic scaling. Research has also indicated that the lack of data load prediction, which aims at the identification of the expected data load increase on peak hours/days, is the root cause of the aforementioned issue. Hence, this paper proposes an enhanced solution, namely, KORDI (Knowledge-based Orchestrated Resource Distribution), aiming at optimising the allocation of Spark resources on Streaming applications in real time with the use of SARIMAX model. The experimental evaluation proves that the proposed solution provides a cost reduction of 38% without affecting stability.
Athanasios Kordelas, Thanasis Spyrou, Spyros Voulgaris, Vasileios Megalooikonomou, Nikos Deligiannis
ISPASS5
2023 Evaluating Quality of Visual Explanations of Deep Learning Models for Vision Tasks
abstract
Explainable artificial intelligence (XAI) has gained considerable attention in recent years as it aims to help humans better understand machine learning decisions, making complex black-box systems more trustworthy. Visual explanation algorithms have been designed to generate heatmaps highlighting image regions that a deep neural network focuses on to make decisions. While convolutional neural network (CNN) models typically follow similar processing operations for feature encoding, the emergence of vision transformer (ViT) has introduced a new approach to machine vision decision-making. Therefore, an important question is which architecture provides more human-understandable explanations. This paper examines the explain-ability of deep architectures, including CNN and ViT models under different vision tasks. To this end, we first performed a subjective experiment asking humans to highlight the key visual features in images that helped them to make decisions in two different vision tasks. Next, using the human-annotated images, ground-truth heatmaps were generated that were compared against heatmaps generated by explanation methods for the deep architectures. Moreover, perturbation tests were performed for objective evaluation of the deep models' explanation heatmaps. According to the results, the explanations generated from ViT are deemed more trustworthy than those produced by other CNNs, and as the features of the input image are more dispersed, the advantage of the model becomes more evident.
Saeed Mahmoudpour, Peter Schelkens, Nikos Deligiannis
QoMEX4
2023 Traffic event detection as a slot filling problem
Giannis Bekoulis, Nikos Deligiannis
Eng. Appl. Artif. Intell.3
2022 NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks
abstract
Natural language explanation (NLE) models aim at explaining the decision-making process of a black box system via generating natural language sentences which are human-friendly, high-level and fine-grained. Current NLE models11Throughout this paper, we refer to NLE models as Natural Language Explanation models aimed for vision and vision-language tasks. explain the decision-making process of a vision or vision-language model (a.k.a., task model), e.g., a VQA model, via a language model (a.k.a., explanation model), e.g., GPT. Other than the additional memory resources and inference time required by the task model, the task and explanation models are completely independent, which disassociates the explanation from the reasoning process made to predict the answer. We introduce NLX-GPT, a general, compact and faithful language model that can simultaneously predict an answer and explain it. We first conduct pre-training on large scale data of image-caption pairs for general understanding of images, and then formulate the answer as a text prediction task along with the explanation. Without region proposals nor a task model, our resulting overall framework attains better evaluation scores, contains much less parameters and is 15× faster than the current SoA model. We then address the problem of evaluating the explanations which can be in many times generic, data-biased and can come in several forms. We therefore design 2 new evaluation measures: (1) explain-predict and (2) retrieval-based attack, a selfevaluation framework that requires no labels. Code is at: https://github.com/fawazsammani/nlxgpt.
Fawaz Sammani, Tanmoy Mukherjee, Nikos Deligiannis
CVPR3
2022 Gradient Variance Loss for Structure-Enhanced Image Super-Resolution
abstract
Recent success in the field of single image super-resolution (SISR) is achieved by optimizing deep convolutional neural networks (CNNs) in the image space with the L1 or L2 loss. However, when trained with these loss functions, models usually fail to recover sharp edges present in the high-resolution (HR) images for the reason that the model tends to give a statistical average of potential HR solutions. During our research, we observe that gradient maps of images generated by the models trained with the L1 or L2 loss have significantly lower variance than the gradient maps of the original high-resolution images. In this work, we propose to alleviate the above issue by introducing a structure-enhancing loss function, coined Gradient Variance (GV) loss, and generate textures with perceptual-pleasant details. Specifically, during the training of the model, we extract patches from the gradient maps of the target and generated output, calculate the variance of each patch and form variance maps for these two images. Further, we minimize the distance between the computed variance maps to enforce the model to produce high variance gradient maps that will lead to the generation of high-resolution images with sharper edges. Experimental results show that the GV loss can significantly improve both Structure Similarity (SSIM) and peak signal-to-noise ratio (PSNR) performance of existing image super-resolution (SR) deep learning models.
Lusine Abrahamyan, Anh Minh Truong, Wilfried Philips, Nikos Deligiannis
ICASSP4
2022 Novel Class Discovery: A Dependency Approach
abstract
Supervised and semi-supervised algorithms have been designed under a closed-world setting, with the assumption that unlabeled data consists of classes previously seen in labeled training data. However, real world is inherently open set where this assumption is often violated, and thus novel data may be encountered in test data. In this paper, we look at the problem where the model is required to discover novel classes never encountered in the labeled set. We propose a dependency measure based on Squared Mutual Information (SMI) where we simultaneously learn to classify and cluster the data. Our experiments show that our approach is able to achieve competitive performance on CIFAR and Imagenet datasets.
Tanmoy Mukherjee, Nikos Deligiannis
ICASSP2
2022 Entropy-Based Feature Extraction for Real-Time Semantic Segmentation
abstract
This paper introduces an efficient patch-based computational module, coined Entropy-based Patch Encoder (EPE) module, for resource-constrained semantic segmentation. The EPE module consists of three lightweight fully-convolutional encoders, each extracting features from image patches with a different amount of entropy. Patches with high entropy are being processed by the encoder with the largest number of parameters, patches with moderate entropy are processed by the encoder with a moderate number of parameters, and patches with low entropy are processed by the smallest encoder. The intuition behind the module is the following: as patches with high entropy contain more information, they need an encoder with more parameters, unlike low entropy patches, which can be processed using a small encoder. Consequently, processing part of the patches via the smaller encoder can significantly reduce the computational cost of the module. Experiments show that EPE can boost the performance of existing real-time semantic segmentation models with a slight increase in the computational cost. Specifically, EPE increases the mIOU performance of DFANet A by 0.9% with only 1.2% increase in the number of parameters and the mIOU performance of EDANet by 1% with 10% increase of the model parameters.
Lusine Abrahamyan, Nikos Deligiannis
ICIP2
2022 Designing CNNs for Multimodal Image Restoration and Fusion via Unfolding the Method of Multipliers
abstract
Multimodal, alias, guided, image restoration is the reconstruction of a degraded image from a target modality with the aid of a high quality image from another modality. A similar task is image fusion; it refers to merging images from different modalities into a composite image. Traditional approaches for multimodal image restoration and fusion include analytical methods that are computationally expensive at inference time. Recently developed deep learning methods have shown a great performance at a reduced computational cost; however, since these methods do not incorporate prior knowledge about the problem at hand, they result in a “black box” model, that is, one can hardly say what the model has learned. In this paper, we formulate multimodal image restoration and fusion as a coupled convolutional sparse coding problem, and adopt the Method of Multipliers (MM) for its solution. Then, we use the MM-based solution to design a convolutional neural network (CNN) encoder that follows the principle of deep unfolding. To address multimodal image restoration and fusion, we design two multimodal models which employ the proposed encoder followed by an appropriately designed decoder that maps the learned representations to the desired output. Unlike most existing deep learning designs comprising multiple encoding branches followed by a concatenation or a linear combination fusion block, the proposed design provides an efficient and structured way to fuse information at different stages of the network, providing representations that can lead to accurate image reconstruction. The proposed models are applied to three image restoration tasks, as well as two image fusion tasks. Quantitative and qualitative comparisons against various state-of-the-art analytical and deep learning methods corroborate the superior performance of the proposed framework.
Iman Marivani, Evaggelia Tsiligianni, Bruno Cornelis, Nikos Deligiannis
IEEE Trans. Circuits Syst. Video Technol.4
2022 Learned Gradient Compression for Distributed Deep Learning
abstract
Training deep neural networks on large datasets containing high-dimensional data requires a large amount of computation. A solution to this problem is data-parallel distributed training, where a model is replicated into several computational nodes that have access to different chunks of the data. This approach, however, entails high communication rates and latency because of the computed gradients that need to be shared among nodes at every iteration. The problem becomes more pronounced in the case that there is wireless communication between the nodes (i.e., due to the limited network bandwidth). To address this problem, various compression methods have been proposed, including sparsification, quantization, and entropy encoding of the gradients. Existing methods leverage the intra-node information redundancy, that is, they compress gradients at each node independently. In contrast, we advocate that the gradients across the nodes are correlated and propose methods to leverage this inter-node redundancy to improve compression efficiency. Depending on the node communication protocol (parameter server or ring-allreduce), we propose two instances for the gradient compression that we coin Learned Gradient Compression (LGC). Our methods exploit an autoencoder (i.e., trained during the first stages of the distributed training) to capture the common information that exists in the gradients of the distributed nodes. To constrain the nodes' computational complexity, the autoencoder is realized with a lightweight neural network. We have tested our LGC methods on the image classification and semantic segmentation tasks using different convolutional neural networks (CNNs) [ResNet50, ResNet101, and pyramid scene parsing network (PSPNet)] and multiple datasets (ImageNet, Cifar10, and CamVid). The ResNet101 model trained for image classification on Cifar10 achieved significant compression rate reductions with the accuracy of 93.57%, which is lower than the baseline distributed training with uncompressed gradients only by 0.18%. The rate of the model is reduced by 8095× and 8× compared with the baseline and the state-of-the-art deep gradient compression (DGC) method, respectively.
Lusine Abrahamyan, Giannis Bekoulis, Nikos Deligiannis
IEEE Trans. Neural Networks Learn. Syst.4
2021 HCGM-Net: A Deep Unfolding Network for Financial Index Tracking
abstract
Tracking the performance of a financial index by selecting a subset of assets composing the index is a problem that raises several difficulties due to the large size of the stock market. Typically, optimisation algorithms with high complexity are employed to address such problems. In this paper, we focus on sparse index tracking and employ a Frank-Wolfe-based algorithm which we translate into a deep neural network, a strategy known as deep unfolding. Numerical experiments show that the learned model outperforms the iterative algorithm, leading to high accuracy at a low computational cost. To the best of our knowledge, this is the first deep unfolding design proposed for financial data processing.
Ruben Pauwels, Evaggelia Tsiligianni, Nikos Deligiannis
ICASSP3
2021 Bias Loss for Mobile Neural Networks
abstract
Compact convolutional neural networks (CNNs) have witnessed exceptional improvements in performance in recent years. However, they still fail to provide the same predictive power as CNNs with a large number of parameters. The diverse and even abundant features captured by the layers is an important characteristic of these successful CNNs. However, differences in this characteristic between large CNNs and their compact counterparts have rarely been investigated. In compact CNNs, due to the limited number of parameters, abundant features are unlikely to be obtained, and feature diversity becomes an essential characteristic. Diverse features present in the activation maps derived from a data point during model inference may indicate the presence of a set of unique descriptors necessary to distinguish between objects of different classes. In contrast, data points with low feature diversity may not provide a sufficient amount of unique descriptors to make a valid prediction; we refer to them as random predictions. Random predictions can negatively impact the optimization process and harm the final performance. This paper proposes addressing the problem raised by random predictions by reshaping the standard cross-entropy to make it biased toward data points with a limited number of unique descriptive features. Our novel Bias Loss focuses the training on a set of valuable data points and prevents the vast number of samples with poor learning features from misleading the optimization process. Furthermore, to show the importance of diversity, we present a family of SkipblockNet models whose architectures are brought to boost the number of unique descriptors in the last layers. Experiments conducted on benchmark datasets demonstrate the superiority of the proposed loss function over the cross-entropy loss. Moreover, our SkipblockNet-M can achieve 1% higher classification accuracy than MobileNetV3 Large with similar computational cost on the ImageNet ILSVRC-2012 classification dataset. The code is available on the link - https://github.com/lusinlu/biasloss_skipblocknet.
Lusine Abrahamyan, Valentin Ziatchin, Nikos Deligiannis
ICCV4
2021 Keynote Lecture : Gradient compression for efficient distributed deep learning
abstract
Summary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Recent successful results in the field of artificial intelligence and machine learning are achieved with deep learning models that contain a large number of parameters and are trained using a massive amount of data. Training such deep networks in a single machine (given a fixed set of hyperparameters) can take weeks. An answer to this problem is data-parallel distributed training, where a deep model is replicated into several computational nodes that have access to different chunks of the data. This approach, however, entails high communication rates and latency because of the computed gradients that need to be shared among nodes at every iteration. We will elaborate on various gradient compression strategies proposed to address this bottleneck within distributed training, including gradient sparsification, quantization, and entropy encoding. We will also discuss error correction techniques that compensate for the errors introduced by gradient compression. Furthermore, we will present new communication strategies that explore the correlation of gradients across distributed nodes to achieve further improvements in reducing the communication rate and latency.
Nikos Deligiannis
ISPDC1
2021 Interpretable Deep Learning for Multimodal Super-Resolution of Medical Images
Evaggelia Tsiligianni, Matina Zerva, Iman Marivani, Nikos Deligiannis, Lisimachos P. Kondi
MICCAI (6)4
2021 Generalization error bounds for deep unfolding RNNs
abstract
Recurrent Neural Networks (RNNs) are powerful models with the ability to model sequential data. However, they are often viewed as black-boxes and lack in interpretability. Deep unfolding methods take a step towards interpretability by designing deep neural networks as learned variations of iterative optimization algorithms to solve various signal processing tasks. In this paper, we explore theoretical aspects of deep unfolding RNNs in terms of their generalization ability. Specifically, we derive generalization error bounds for a class of deep unfolding RNNs via Rademacher complexity analysis. To our knowledge, these are the first generalization bounds proposed for deep unfolding RNNs. We show theoretically that our bounds are tighter than similar ones for other recent RNNs, in terms of the number of timesteps. By training models in a classification setting, we demonstrate that deep unfolding RNNs can outperform traditional RNNs in standard sequence classification tasks. These experiments allow us to relate the empirical generalization error to the theoretical bounds. In particular, we show that over-parametrized deep unfolding models like reweighted-RNN achieve tight theoretical error bounds with minimal decrease in accuracy, when trained with explicit regularization.
Boris Joukovsky, Tanmoy Mukherjee, Huynh Van Luong, Nikos Deligiannis
UAI4
2021 Graph convolutional neural networks with node transition probability-based message passing and DropNode regularization
Tien Do Huu, Duc Minh Nguyen 0002, Giannis Bekoulis, Adrian Munteanu 0001, Nikos Deligiannis
Expert Syst. Appl.5
2021 Designing Interpretable Recurrent Neural Networks for Video Reconstruction via Deep Unfolding
abstract
Deep unfolding methods design deep neural networks as learned variations of optimization algorithms through the unrolling of their iterations. These networks have been shown to achieve faster convergence and higher accuracy than the original optimization methods. In this line of research, this paper presents novel interpretable deep recurrent neural networks (RNNs), designed by the unfolding of iterative algorithms that solve the task of sequential signal reconstruction (in particular, video reconstruction). The proposed networks are designed by accounting that video frames’ patches have a sparse representation and the temporal difference between consecutive representations is also sparse. Specifically, we design an interpretable deep RNN (coined reweighted-RNN) by unrolling the iterations of a proximal method that solves a reweighted version of the$\ell _{1}$-$\ell _{1}$minimization problem. Due to the underlying minimization model, our reweighted-RNN has a different thresholding function (alias, different activation function) for each hidden unit in each layer. In this way, it has higher network expressivity than existing deep unfolding RNN models. We also present the derivative$\ell _{1}$-$\ell _{1}$-RNN model, which is obtained by unfolding a proximal method for the$\ell _{1}$-$\ell _{1}$minimization problem. We apply the proposed interpretable RNNs to the task of video frame reconstruction from low-dimensional measurements, that is, sequential video frame reconstruction. The experimental results on various datasets demonstrate that the proposed deep RNNs outperform various RNN models.
Huynh Van Luong, Boris Joukovsky, Nikos Deligiannis
IEEE Trans. Image Process.3
2020 Graph Auto-Encoder for Graph Signal Denoising
abstract
Signal denoising is an important problem with a vast literature. Recently, signal denoising on graphs has received a lot of attention due to the increasing use of graph-structured signals. However, well-established signal denoising methods do not generalize to graph signals with irregular structures, while existing graph denoising methods do not capture well the abstract representations inherent in the signals. To bridge this gap, we propose to use graph convolutional neural network with a Kron-reduction-based pooling operator for denoising on graphs. The proposed model can effectively capture the irregular data structure and learn the underlying representations in the signals, leading to improved performance over existing methods in experiments involving real-world traffic signals.
Tien Do Huu, Duc Minh Nguyen 0002, Nikos Deligiannis
ICASSP3
2020 Joint Image Super-Resolution Via Recurrent Convolutional Neural Networks With Coupled Sparse Priors
abstract
Joint image super-resolution (SR) refers to the reconstruction of a high-resolution image from its low-resolution version with the aid of a high-resolution image from another modality. Inspired by the recent success of recurrent neural networks in single image SR, we propose a novel multimodal recurrent convolutional neural network with coupled sparse priors for joint image SR. Our network fuses representations of the two image modalities at input layers using a learned multimodal convolutional sparse coding network. Additional recurrent convolutional stages are performed to further learn the mapping between the input modalities and the desired highresolution estimate. We apply the proposed network to the tasks of near-infrared image SR and multi-spectral image SR using RGB images as the guidance modality. Experimental results show the superior performance of the proposed multimodal recurrent convolutional network against several state-of-the-art single-modal and multimodal image SR methods.
Iman Marivani, Evaggelia Tsiligianni, Bruno Cornelis, Nikos Deligiannis
ICIP4
2020 Temporal Collaborative Filtering with Graph Convolutional Neural Networks
abstract
Temporal collaborative filtering (TCF) methods aim at modelling non-static aspects behind recommender systems, such as the dynamics in users' preferences and social trends around items. State-of-the-art TCF methods employ recurrent neural networks (RNNs) to model such aspects. These methods deploy matrix-factorization-based (MF -based) approaches to learn the user and item representations. Recently, graph-neural-network-based (GNN-based) approaches have shown improved performance in providing accurate recommendations over traditional MF -based approaches in non-temporal CF settings. Motivated by this, we propose a novel TCF method that leverages GNNs to learn user and item representations, and RNNs to model their temporal dynamics. A challenge with this method lies in the increased data sparsity, which negatively impacts obtaining meaningful quality representations with GNNs. To overcome this challenge, we train a GNN model at each time step using a set of observed interactions accumulated time-wise. Comprehensive experiments on real-world data show the improved performance obtained by our method over several state-of-the-art temporal and non-temporal CF models.
Esther Rodrigo Bonet, Duc Minh Nguyen 0002, Nikos Deligiannis
ICPR3
2020 Graph-Deep-Learning-Based Inference of Fine-Grained Air Quality From Mobile IoT Sensors
abstract
Internet-of-Things (IoT) technologies incorporate a large number of different sensing devices and communication technologies to collect a large amount of data for various applications. Smart cities employ IoT infrastructures to build services useful for the administration of the city and the citizens. In this article, we present an IoT pipeline for acquisition, processing, and visualization of air pollution data over the city of Antwerp, Belgium. Our system employs IoT devices mounted on vehicles as well as static reference stations to measure a variety of city parameters, such as humidity, temperature, and air pollution. Mobile measurements cover a larger area compared to static stations; however, there is a tradeoff between temporal and spatial resolution. We address this problem as a matrix completion on graphs problem and rely on variational graph autoencoders to propose a deep learning solution for the estimation of the unknown air pollution values. Our model is extended to capture the correlation among different air pollutants, leading to improved estimation. We conduct experiments at different spatial and temporal resolution and compare with state-of-the-art methods to show the efficiency of our approach. The observed and estimated air pollution values can be accessed by interested users through a Web visualization tool designed to provide an air pollution map of the city of Antwerp.
Tien Do Huu, Evaggelia Tsiligianni, Xuening Qin, Jelle Hofman, Valerio Panzica La Manna, Wilfried Philips, Nikos Deligiannis
IEEE Internet Things J.7
2020 DeepFPC: A deep unfolded network for sparse signal recovery from 1-Bit measurements with application to DOA estimation
Peng Xiao 0005, Bin Liao 0001, Nikos Deligiannis
Signal Process.3
2020 Multimodal Deep Unfolding for Guided Image Super-Resolution
abstract
The reconstruction of a high resolution image given a low resolution observation is an ill-posed inverse problem in imaging. Deep learning methods rely on training data to learn an end-to-end mapping from a low-resolution input to a high-resolution output. Unlike existing deep multimodal models that do not incorporate domain knowledge about the problem, we propose a multimodal deep learning design that incorporates sparse priors and allows the effective integration of information from another image modality into the network architecture. Our solution relies on a novel deep unfolding operator, performing steps similar to an iterative algorithm for convolutional sparse coding with side information; therefore, the proposed neural network is interpretable by design. The deep unfolding architecture is used as a core component of a multimodal framework for guided image super-resolution. An alternative multimodal design is investigated by employing residual learning to improve the training efficiency. The presented multimodal approach is applied to super-resolution of near-infrared and multi-spectral images as well as depth upsampling using RGB images as side information. Experimental results show that our model outperforms state-of-the-art methods.
Iman Marivani, Evaggelia Tsiligianni, Bruno Cornelis, Nikos Deligiannis
IEEE Trans. Image Process.4
2020 Geometric Matrix Completion With Deep Conditional Random Fields
abstract
The problem of completing high-dimensional matrices from a limited set of observations arises in many big data applications, especially recommender systems. The existing matrix completion models generally follow either a memory- or a model-based approach, whereas geometric matrix completion (GMC) models combine the best from both approaches. Existing deep-learning-based geometric models yield good performance, but, in order to operate, they require a fixed structure graph capturing the relationships among the users and items. This graph is typically constructed by evaluating a pre-defined similarity metric on the available observations or by using side information, e.g., user profiles. In contrast, Markov-random-fields-based models do not require a fixed structure graph but rely on handcrafted features to make predictions. When no side information is available and the number of available observations becomes very low, existing solutions are pushed to their limits. In this article, we propose a GMC approach that addresses these challenges. We consider matrix completion as a structured prediction problem in a conditional random field (CRF), which is characterized by a maximum a posteriori (MAP) inference, and we propose a deep model that predicts the missing entries by solving the MAP inference problem. The proposed model simultaneously learns the similarities among matrix entries, computes the CRF potentials, and solves the inference problem. Its training is performed in an end-to-end manner, with a method to supervise the learning of entry similarities. Comprehensive experiments demonstrate the superior performance of the proposed model compared to various state-of-the-art models on popular benchmark data sets and underline its superior capacity to deal with highly incomplete matrices.
Duc Minh Nguyen 0002, A. Robert Calderbank, Nikos Deligiannis
IEEE Trans. Neural Networks Learn. Syst.3
2019 Matrix Completion with Variational Graph Autoencoders: Application in Hyperlocal Air Quality Inference
abstract
Inferring air quality from a limited number of observations is an essential task for monitoring and controlling air pollution. Existing inference methods typically use low spatial resolution data collected by fixed monitoring stations and infer the concentration of air pollutants using additional types of data, e.g., meteorological and traffic information. In this work, we focus on street-level air quality inference by utilizing data collected by mobile stations. We formulate air quality inference in this setting as a graph-based matrix completion problem and propose a novel variational model based on graph convolutional autoencoders. Our model captures effectively the spatio-temporal correlation of the measurements and does not depend on the availability of additional information apart from the street-network topology. Experiments on a real air quality dataset, collected with mobile stations, shows that the proposed model outperforms state-of-the-art approaches.
Tien Do Huu, Duc Minh Nguyen 0002, Evaggelia Tsiligianni, Angel Lopez Aguirre, Valerio Panzica La Manna, Frank J. Pasveer, Wilfried Philips, Nikos Deligiannis
ICASSP8
2019 Designing Recurrent Neural Networks by Unfolding an L1-L1 Minimization Algorithm
abstract
We propose a new deep recurrent neural network (RNN) architecture for sequential signal reconstruction. Our network is designed by unfolding the iterations of the proximal gradient method that solves the ℓ1-ℓ1minimization problem. As such, our network leverages by design that signals have a sparse representation and that the difference between consecutive signal representations is also sparse. We evaluate the proposed model in the task of reconstructing video frames from compressive measurements and show that it outperforms several state-of-the-art RNN models.
Hung Duy Le, Huynh Van Luong, Nikos Deligiannis
ICIP3
2019 Learned Multimodal Convolutional Sparse Coding for Guided Image Super-Resolution
abstract
The success of deep learning in various tasks, including solving inverse problems, has triggered the need for designing deep neural networks that incorporate domain knowledge. In this paper, we design a multimodal deep learning architecture for guided image super-resolution, which refers to the problem of super-resolving a low-resolution image with the aid of a high-resolution image of another modality. The proposed architecture is based on a novel deep learning model, obtained by unfolding a proximal method that solves the problem of convolutional sparse coding with side information. We applied the proposed architecture to super-resolve near-infrared images using RGB images as side information. Experimental results report average PSNR gains of up to 2.85 dB against state-of-the-art multimodal deep learning and sparse coding models.
Iman Marivani, Evaggelia Tsiligianni, Bruno Cornelis, Nikos Deligiannis
ICIP4
2019 Deep Coupled-Representation Learning for Sparse Linear Inverse Problems With Side Information
abstract
In linear inverse problems, the goal is to recover a target signal from undersampled, incomplete or noisy linear measurements. Typically, the recovery relies on complex numerical optimization methods; recent approaches perform an unfolding of a numerical algorithm into a neural network form, resulting in a substantial reduction of the computational complexity. In this letter, we consider the recovery of a target signal with the aid of a correlated signal, the so-called side information (SI), and propose a deep unfolding model that incorporates SI. The proposed model is used to learn coupled representations of correlated signals from different modalities, enabling the recovery of multi-modal data at a low computational cost. As such, our work introduces the first deep unfolding method with SI, which actually comes from a different modality. We apply our model to reconstruct near-infrared images from undersampled measurements given RGB images as SI. Experimental results demonstrate the superior performance of the proposed framework against single-modal deep learning methods that do not use SI, multi-modal deep learning designs, and optimization algorithms.
Evaggelia Tsiligianni, Nikos Deligiannis
IEEE Signal Process. Lett.2
2018 Online Decomposition of Compressive Streaming Data Using n-l1 Cluster-Weighted Minimization
abstract
We consider a decomposition method for compressive streaming data in the context of online compressive Robust Principle Component Analysis (RPCA). The proposed decomposition solves an n-ℓ1 cluster-weighted minimization to decompose a sequence of frames (or vectors), into sparse and low-rank components from compressive measurements. Our method processes a data vector of the stream per time instance from a small number of measurements in contrast to conventional batch RPCA, which needs to access full data. The n-ℓ1 cluster-weighted minimization leverages the sparse components along with their correlations with multiple previously-recovered sparse vectors. Moreover, the proposed minimization can exploit the structures of sparse components via clustering and re-weighting iteratively. The method outperforms the existing methods for both numerical data and actual video data.
Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup
DCC2
2018 Twitter User Geolocation Using Deep Multiview Learning
abstract
Predicting the geographical location of users on social networks like Twitter is an active research topic with plenty of methods proposed so far. Most of the existing work follows either a content-based or a network-based approach. The former is based on user-generated content while the latter exploits the structure of the network of users. In this paper, we propose a more generic approach, which incorporates not only both content-based and network-based features, but also other available information into a unified model. Our approach, named Multi-Entry Neural Network (MENET), leverages the latest advances in deep learning and multiview learning. A realization of MENET with textual, network and metadata features results in an effective method for Twitter user geolocation, achieving the state of the art on two well-known datasets.
Tien Do Huu, Duc Minh Nguyen 0002, Evaggelia Tsiligianni, Bruno Cornelis, Nikos Deligiannis
ICASSP5
2018 Extendable Neural Matrix Completion
abstract
Matrix completion is one of the key problems in signal processing and machine learning, with applications ranging from image processing and data gathering to classification and recommender systems. Recently, deep neural networks have been proposed as latent factor models for matrix completion and have achieved state-of-the-art performance. Nevertheless, a major problem with existing neural-network-based models is their limited capabilities to extend to samples unavailable at the training stage. In this paper, we propose a deep two-branch neural network model for matrix completion. The proposed model not only inherits the predictive power of neural networks, but is also capable of extending to partially observed samples outside the training set, without the need of retraining or fine-tuning. Experimental studies on popular movie rating datasets prove the effectiveness of our model compared to the state of the art, in terms of both accuracy and extendability.
Duc Minh Nguyen 0002, Evaggelia Tsiligianni, Nikos Deligiannis
ICASSP3
2018 Data aggregation and recovery for the Internet of Things: A compressive demixing approach
abstract
Large-scale wireless sensor networks (WSNs) and Internet-of-Things (IoT) applications involve diverse sensing devices collecting and transmitting massive amounts of heterogeneous data. In this paper, we propose a novel compressive data aggregation and recovery mechanism that reduces the global communication cost without introducing computational overhead at the network nodes. Following the principles of compressive demixing, each node of the network collects measurement readings from multiple sources and mixes them with readings from other nodes into a single low-dimensional measurement vector, which is then relayed to other nodes; the constituent signals are recovered at the sink using convex optimization. Our design achieves significant reduction in the overall network data rates compared to prior schemes based on (distributed) compressed sensing or compressed sensing with (multiple) side information. Experiments using real large-scale air-quality data demonstrate the superior performance of the proposed framework against state-of-the-art solutions, with and without the presence of measurement and transmission noise.
Evangelos Zimos, João F. C. Mota, Evaggelia Tsiligianni, Miguel R. D. Rodrigues, Nikos Deligiannis
WCNC5
2018 Sparse signal recovery with multiple prior information: Algorithm and measurement bounds
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup
Signal Process.2
2018 Learning Discrete Matrix Factorization Models
abstract
Matrix factorization is among the most popular approaches for matrix completion, with recent advances including gradient-based and deep-learning-based methods. Even though many applications involve matrices with discrete values, most of the existing matrix factorization models focus on the continuous domain. Discretization is applied as an additional step, often using a heuristic mapping that results in sub-optimal solutions, which either do not take into account the structure of the matrix or introduce significant quantization errors. In this letter, we propose a novel method that allows gradient-based and deep-learning-based methods to jointly learn both the matrix factorization model and a discretization operator. By introducing a loss function that accounts for the reconstruction error with respect to the discrete predictions, we obtain a discrete matrix completion algorithm with high reconstruction accuracy. Experiments using well-known datasets show the improvement obtained by the proposed algorithm over the state of the art.
Duc Minh Nguyen 0002, Evaggelia Tsiligianni, Nikos Deligiannis
IEEE Signal Process. Lett.3
2018 Compressive Online Robust Principal Component Analysis via n-ℓ1 Minimization
abstract
This paper considers online robust principal component analysis (RPCA) in time-varying decomposition problems such as video foreground-background separation. We propose a compressive online RPCA algorithm that decomposes recursively a sequence of data vectors (e.g., frames) into sparse and low-rank components. Different from conventional batch RPCA, which processes all the data directly, our approach considers a small set of measurements taken per data vector (frame). Moreover, our algorithm can incorporate multiple prior information from previous decomposed vectors via proposing an - minimization method. At each time instance, the algorithm recovers the sparse vector by solving the - minimization problem-which promotes not only the sparsity of the vector but also its correlation with multiple previously recovered sparse vectors-and, subsequently, updates the low-rank component using incremental singular value decomposition. We also establish theoretical bounds on the number of measurements required to guarantee successful compressive separation under the assumptions of static or slowly changing low-rank components. We evaluate the proposed algorithm using numerical experiments and online video foreground-background separation experiments. The experimental results show that the proposed method outperforms the existing methods.
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup
IEEE Trans. Image Process.2
2017 Rate-distortion trade-offs in acquisition of signal parameters
abstract
We consider problems where one wishes to represent a parameter associated with a signal source - subject to a certain rate and distortion - based on the observation of a number of realizations of the source signal. By reducing these indirect vector quantization problems to a standard vector quantization one, we provide a bound to the fundamental interplay between the rate and distortion in the large-rate setting. We specialize this characterization to two particular quantization scenarios: i) the representation of the mean of a multivariate Gaussian source; and ii) the representation of the eigen-spectrum of a multivariate Gaussian source. Numerical results compare our quantization approach to an approach where one recovers the parameters from the representation of the source signals itself: in addition to revealing that the characterization is sharp in the large-rate setting, the results also show that our approach offers considerable gains.
Miguel R. D. Rodrigues, Nikos Deligiannis, Lifeng Lai, Yonina C. Eldar
ICASSP2
2017 Heterogeneous Networked Data Recovery From Compressive Measurements Using a Copula Prior
abstract
Large-scale data collection by means of wireless sensor network and Internet-of-Things technology poses various challenges in view of the limitations in transmission, computation, and energy resources of the associated wireless devices. Compressive data gathering based on compressed sensing has been proven a well-suited solution to the problem. Existing designs exploit the spatiotemporal correlations among data collected by a specific sensing modality. However, many applications, such as environmental monitoring, involve collecting heterogeneous data that are intrinsically correlated. In this paper, we propose to leverage the correlation from multiple heterogeneous signals when recovering the data from compressive measurements. To this end, we propose a novel recovery algorithm-built upon belief-propagation principles-that leverages correlated information from multiple heterogeneous signals. To efficiently capture the statistical dependencies among diverse sensor data, the proposed algorithm uses the statistical model of copula functions. Experiments with heterogeneous air-pollution sensor measurements show that the proposed design provides significant performance improvements against the state-of-the-art compressive data gathering and recovery schemes that use classical compressed sensing, compressed sensing with side information, and distributed compressed sensing.
Nikos Deligiannis, João F. C. Mota, Evangelos Zimos, Miguel R. D. Rodrigues
IEEE Trans. Commun.1
2017 Multi-Modal Dictionary Learning for Image Separation With Application in Art Investigation
abstract
In support of art investigation, we propose a new source separation method that unmixes a single X-ray scan acquired from double-sided paintings. In this problem, the X-ray signals to be separated have similar morphological characteristics, which brings previous source separation methods to their limits. Our solution is to use photographs taken from the front-and back-side of the panel to drive the separation process. The crux of our approach relies on the coupling of the two imaging modalities (photographs and X-rays) using a novel coupled dictionary learning framework able to capture both common and disparate features across the modalities using parsimonious representations; the common component captures features shared by the multi-modal images, whereas the innovation component captures modality-specific information. As such, our model enables the formulation of appropriately regularized convex optimization procedures that lead to the accurate separation of the X-rays. Our dictionary learning framework can be tailored both to a single- and a multi-scale framework, with the latter leading to a significant performance improvement. Moreover, to improve further on the visual quality of the separated images, we propose to train coupled dictionaries that ignore certain parts of the painting corresponding to craquelure. Experimentation on synthetic and real data - taken from digital acquisition of the Ghent Altarpiece (1432) - confirms the superiority of our method against the state-of-the-art morphological component analysis technique that uses either fixed or trained dictionaries to perform image separation.
Nikos Deligiannis, João F. C. Mota, Bruno Cornelis, Miguel R. D. Rodrigues, Ingrid Daubechies
IEEE Trans. Image Process.1
2017 Compressed Sensing With Prior Information: Strategies, Geometry, and Bounds
abstract
We address the problem of compressed sensing (CS) with prior information: reconstruct a target CS signal with the aid of a similar signal that is known beforehand, our prior information. We integrate the additional knowledge of the similar signal into CS via l1-l1and l1-l2minimization. We then establish bounds on the number of measurements required by these problems to successfully reconstruct the original signal. Our bounds and geometrical interpretations reveal that if the prior information has good enough quality, l1-l1minimization improves the performance of CS dramatically. In contrast, l1-l2minimization has a performance very similar to classical CS, and brings no significant benefits. In addition, we use the insight provided by our bounds to design practical schemes to improve prior information. All our findings are illustrated with experimental results.
João F. C. Mota, Nikos Deligiannis, Miguel R. D. Rodrigues
IEEE Trans. Inf. Theory2
2016 The Rate Loss in Binary Source Coding with Decoder Side Information
abstract
Summary form only given. Motivated by the correlation channel modeling problem in practical applications, such as distributed video coding, we study the binary source coding of a uniform source with side information, under asymmetric correlation channel assumptions. First, we consider the case where side information is available to both the encoder and decoder, and give an analytical formula for the rate-distortion bound. Then, we consider the side information to be available only to the decoder and present the derivation of the associated Wyner-Ziv rate-distortion bound. Most importantly, we characterize the evolution of the rate-loss suffered by Wyner-Ziv coding, for all possible binary asymmetric correlation channels.
Andrei Sechelea, Adrian Munteanu 0001, Samuel Cheng 0001, Nikos Deligiannis
DCC4
2016 Bayesian Compressed Sensing with Heterogeneous Side Information
abstract
The classical compressed sensing (CS) paradigm can be modified so as to leverage a signal correlated to the signal of interest, called side information, which is assumed to be provided a priori at the decoder in order to aid reconstruction. In this work, we propose a novel CS reconstruction method based on belief propagation principles, which manages to exploit side information generated from a diverse (or heterogeneous) data source by using the statistical model of copula functions. Through simulations, we demonstrate that the proposed method yields significant reduction in the mean-squared error of the reconstructed signal as compared to state-of-the-art methods in classical compressed sensing and compressed sensing with side information.
Evangelos Zimos, João F. C. Mota, Miguel R. D. Rodrigues, Nikos Deligiannis
DCC4
2016 Reference-based compressed sensing: A sample complexity approach
abstract
We address the problem of reference-based compressed sensing: reconstruct a sparse signal from few linear measurements using as prior information a reference signal, a signal similar to the signal we want to reconstruct. Access to reference signals arises in applications such as medical imaging, e.g., through prior images of the same patient, and compressive video, where previously reconstructed frames can be used as reference. Our goal is to use the reference signal to reduce the number of required measurements for reconstruction. We achieve this via a reweighted ℓ1-ℓ1minimization scheme that updates its weights based on a sample complexity bound. The scheme is simple, intuitive and, as our experiments show, outperforms prior algorithms, including reweighted ℓ1minimization, ℓ1-ℓ1minimization, and modified CS.
João F. C. Mota, Lior Weizman, Nikos Deligiannis, Yonina C. Eldar, Miguel R. D. Rodrigues
ICASSP3
2016 X-ray image separation via coupled dictionary learning
abstract
In support of art investigation, we propose a new source separation method that unmixes a single X-ray scan acquired from double-sided paintings. Unlike prior source separation methods, which are based on statistical or structural incoherence of the sources, we use visual images taken from the front- and back-side of the panel to drive the separation process. The coupling of the two imaging modalities is achieved via a new multi-scale dictionary learning method. Experimental results demonstrate that our method succeeds in the discrimination of the sources, while state-of-the-art methods fail to do so.
Nikos Deligiannis, João F. C. Mota, Bruno Cornelis, Miguel R. D. Rodrigues, Ingrid Daubechies
ICIP1
2016 Distributed coding of multiview sparse sources with joint recovery
abstract
In support of applications involving multiview sources in distributed object recognition using lightweight cameras, we propose a new method for the distributed coding of sparse sources as visual descriptor histograms extracted from multiview images. The problem is challenging due to the computational and energy constraints at each camera as well as the limitations regarding inter-camera communication. Our approach addresses these challenges by exploiting the sparsity of the visual descriptor histograms as well as their intra- and inter-camera correlations. Our method couples distributed source coding of the sparse sources with a new joint recovery algorithm that incorporates multiple side information signals, where prior knowledge (low quality) of all the sparse sources is initially sent to exploit their correlations. Experimental evaluation using the histograms of shift-invariant feature transform (SIFT) descriptors extracted from multiview images shows that our method leads to an average bit-rate saving of 30.7% compared to the state-of-the-art distributed compressed sensing method with independent encoding of the sources.
Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup
PCS2
2016 On the Rate-Distortion Function for Binary Source Coding With Side Information
abstract
We present an in-depth analysis of the problem of lossy compression of binary sources in the presence of correlated side information, where the correlation is given by a generic binary asymmetric channel and the Hamming distance is the distortion metric. Our analysis is motivated by systematic rate-distortion gains observed when applying asymmetric correlation models in Wyner-Ziv video coding. First, we derive for the first time the rate-distortion function for conventional predictive coding in the binary-asymmetric-correlation-channel scenario. Second, we propose a new bound for the case where the side information is only available at the decoder-Wyner-Ziv coding. We conjecture this bound to be tight. We show that the maximum rate needed to encode as well as the maximum rate-loss of Wyner-Ziv coding relative to predictive coding corresponds to uniform sources and symmetric correlations. Importantly, we show that the upper bound on the rate-loss established by Zamir is not tight and that the maximum value is actually significantly lower. Moreover, we prove that the only binary correlation channel that incurs no rate-loss for Wyner-Ziv coding compared with predictive coding is the Z-channel. Finally, we complement our analysis with new compression performance results obtained with our state-of-the-art Wyner-Ziv video coding system.
Andrei Sechelea, Adrian Munteanu 0001, Samuel Cheng 0001, Nikos Deligiannis
IEEE Trans. Commun.4
2015 Dynamic sparse state estimation using ℓ1-ℓ1 minimization: Adaptive-rate measurement bounds, algorithms and applications
abstract
We propose a recursive algorithm for estimating time-varying signals from a few linear measurements. The signals are assumed sparse, with unknown support, and are described by a dynamical model. In each iteration, the algorithm solves an ℓ1-ℓ1minimization problem and estimates the number of measurements that it has to take at the next iteration. These estimates are computed based on recent theoretical results for ℓ1-ℓ1minimization. We also provide sufficient conditions for perfect signal reconstruction at each time instant as a function of an algorithm parameter. The algorithm exhibits high performance in compressive tracking on a real video sequence, as shown in our experimental results.
João F. C. Mota, Nikos Deligiannis, Aswin C. Sankaranarayanan, Volkan Cevher, Miguel R. D. Rodrigues
ICASSP2
2015 Vectors of locally aggregated centers for compact video representation
abstract
We propose a novel vector aggregation technique for compact video representation, with application in accurate similarity detection within large video datasets. The current state-of-the-art in visual search is formed by the vector of locally aggregated descriptors (VLAD) of Jegou et al. VLAD generates compact video representations based on scale-invariant feature transform (SIFT) vectors (extracted per frame) and local feature centers computed over a training set. With the aim to increase robustness to visual distortions, we propose a new approach that operates at a coarser level in the feature representation. We create vectors of locally aggregated centers (VLAC) by first clustering SIFT features to obtain local feature centers (LFCs) and then encoding the latter with respect to given centers of local feature centers (CLFCs), extracted from a training set. The sum-of-differences between the LFCs and the CLFCs are aggregated to generate an extremely-compact video description used for accurate video segment similarity detection. Experimentation using a video dataset, comprising more than 1000 minutes of content from the Open Video Project, shows that VLAC obtains substantial gains in terms of mean Average Precision (mAP) against VLAD and the hyper-pooling method of Douze et al., under the same compaction factor and the same set of distortions.
Alhabib Abbas, Nikos Deligiannis, Yiannis Andreopoulos
ICME2
2015 Decentralized multichannel medium access control: viewing desynchronization as a convex optimization method
abstract
Desynchronization algorithms are essential in the design of collision-free medium access control (MAC) mechanisms for wireless sensor networks. Desync is a well-known desynchronization algorithm that operates under limited listening. In this paper, we view Desync as a gradient descent method solving a convex optimization problem. This enables the design of a novel decentralized, collision-free, multichannel medium access control (MAC) algorithm. Moreover, by using Nesterov's fast gradient method, we obtain a new algorithm that converges to the steady network state much faster. Simulations and experimental results on an IEEE 802.15.4-based wireless sensor network deployment show that our algorithms achieve significantly faster convergence to steady network state and substantially higher throughput compared to the recently standardized IEEE 802.15.4e-2012 time synchronized channel hopping (TSCH) scheme. In addition, our mechanism has a comparable power dissipation with respect to TSCH and does not need a coordinator node or coordination channel.
Nikos Deligiannis, João F. C. Mota, George Smart, Yiannis Andreopoulos
IPSN1
2015 Decentralized time-synchronized channel swapping
abstract
We are working on a new concept for decentralized medium access control (MAC), termed decentralized time-synchronized channel swapping (DT-SCS). Under the proposed DT-SCS and its associated MAC-layer protocol, wireless nodes converge to synchronous beacon packet transmissions across all IEEE802.15.4 channels, with balanced numbers of nodes in each channel. This is achieved by reactive listening mechanisms, based on pulse coupled oscillator techniques. Once convergence to the multichannel time-synchronized state is achieved, peer-to-peer channel swapping can then take place via swap requests and acknowledgments made by concurrent transmitters in neighboring channels. Our implementation of DT-SCS reveals that our proposal comprises an excellent candidate for completely decentralized MAC-layer coordination in WSNs by providing for quick convergence to steady state, high bandwidth utilization, high connectivity and robustness to interference and hidden nodes. The demo will showcase the properties of DT-SCS and will also present its behaviour under various scenarios for hidden nodes and interference, both experimentally and with the help of visualization of simulation results.
George Smart, Nikos Deligiannis, João F. C. Mota, Yiannis Andreopoulos
IPSN2
2015 Fast Desynchronization for Decentralized Multichannel Medium Access Control
abstract
Distributed desynchronization algorithms are key to wireless sensor networks as they allow for medium access control in a decentralized manner. In this paper, we view desynchronization primitives as iterative methods that solve optimization problems. In particular, by formalizing a well established desynchronization algorithm as a gradient descent method, we establish novel upper bounds on the number of iterations required to reach convergence. Moreover, by using Nesterov's accelerated gradient method, we propose a novel desynchronization primitive that provides for faster convergence to the steady state. Importantly, we propose a novel algorithm that leads to decentralized time-synchronous multichannel TDMA coordination by formulating this task as an optimization problem. Our simulations and experiments on a densely-connected IEEE 802.15.4-based wireless sensor network demonstrate that our scheme provides for faster convergence to the steady state, robustness to hidden nodes, higher network throughput and comparable power dissipation with respect to the recently standardized IEEE 802.15.4e-2012 time-synchronized channel hopping (TSCH) scheme.
Nikos Deligiannis, João F. C. Mota, George Smart, Yiannis Andreopoulos
IEEE Trans. Commun.1
2014 On the stochastic modeling of desynchronization convergence in wireless sensor networks
abstract
Desynchronization is a fundamental approach in wireless sensor networks that allows for convergence to time-division multiple access (TDMA) of the medium without the need for clock synchronization and centralized coordination. The method is based on the concept of reactive listening of periodic fre message broadcasts between nodes sharing the given spectrum. We propose a novel framework to estimate the required iterations for convergence to fair TDMA scheduling. Unlike previous conjectures or bounds found in the literature, our estimation framework is based on a stochastic modeling approach. Experiments via imote2 TinyOS nodes and simulations demonstrate that the proposed estimates characterize the experimental desynchronization convergence iterations signifcantly better than existing conjectures or bounds.
Dujdow Buranapanichkit, Nikos Deligiannis, Yiannis Andreopoulos
ICASSP2
2014 Progressively refined wyner-ziv video coding for visual sensors
abstract
Wyner-Ziv video coding constitutes an alluring paradigm for visual sensor networks, offering efficient video compression with low complexity encoding characteristics. This work presents a novel hash-driven Wyner-Ziv video coding architecture for visual sensors, implementing the principles of successively refined Wyner-Ziv coding. To this end, so-called side-information refinement levels are constructed for a number of grouped frequency bands of the discrete cosine transform. The proposed codec creates side-information by means of an original overlapped block motion estimation and pixel-based multihypothesis prediction technique, specifically built around the pursued refinement strategy. The quality of the side-information generated at every refinement level is successively improved, leading to gradually enhanced Wyner-Ziv coding performance. Additionally, this work explores several temporal prediction structures, including a new hierarchical unidirectional prediction structure, providing both temporal scalability and low delay coding. Experimental results include a thorough evaluation of our novel Wyner-Ziv codec, assessing the impact of the proposed successive refinement scheme and the supported temporal prediction structures for a wide range of hash configurations and group of pictures sizes. The results report significant compression gains with respect to benchmark systems in Wyner-Ziv video coding (e.g., up to 42.03% over DISCOVER) as well as versus alternative state-of-the-art schemes refining the side-information.
Nikos Deligiannis, Frederik Verbist, Jürgen Slowack, Rik Van de Walle, Peter Schelkens, Adrian Munteanu 0001
ACM Trans. Sens. Networks1
2013 Genome Sequence Compression with Distributed Source Coding
abstract
In this paper, we develop a novel genome compression framework based on distributed source coding (DSC)[3], which is specially tailored to the need of miniaturized devices. At the encoder side, subsequences with adaptive code length can be compressed flexibly through either low complexity DSC based syndrome coding or hash coding with the decision determined by the existence of variations between source and reference known from the decoder feedback. Moreover, to tackle the variations between source and reference at the decoder, we carefully designed a factor graph based low-density parity-check (LDPC) decoder, which automatically detects insertion, deletion and substitution.
Shuang Wang 0002, Xiaoqian Jiang, Lijuan Cui, Wenrui Dai, Nikos Deligiannis, Pinghao Li, Hongkai Xiong, Samuel Cheng 0001, Lucila Ohno-Machado
DCC5
2013 Optimized segmentation of H.264/AVC video for HTTP adaptive streaming
Jan Lievens, Shahid M. Satti, Nikos Deligiannis, Peter Schelkens, Adrian Munteanu 0001
IM3
2013 Probabilistic motion-compensated prediction in distributed video coding
Frederik Verbist, Nikos Deligiannis, Joeri Barbarien, Peter Schelkens, Adrian Munteanu 0001, Jan Cornelis 0001
Multim. Tools Appl.2
2012 An optimization algorithm for scalable multiple description scalar quantizers
Shahid M. Satti, Nikos Deligiannis, Adrian Munteanu 0001, Peter Schelkens, Jan Cornelis 0001
ISITA2
2012 Feedback-constrained Wyner-Ziv video coding
abstract
Distributed video coding (DVC) systems described in the literature often make use of a feedback channel to determine the rate. However, supporting such a feedback channel in practice may be difficult, particularly considering that current approaches are unable to incorporate constraints on feedback channel usage. Therefore, in this paper we propose a system in which the number of requests per Wyner-Ziv frame can be constrained to a fixed value. Our technique involves decoder-side rate estimation, modeling of its accuracy, and finally defining the number of Wyner-Ziv bits for each of the requests allowed. The experimental results indicate that the performance loss is limited when compared to a configuration operating without constraints on the feedback channel.
Jürgen Slowack, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
PCS2
2012 Decoder-driven mode decision in a block-based distributed video codec
Stefaan Mys, Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Multim. Tools Appl.4
2012 Efficient Low-Delay Distributed Video Coding
abstract
Distributed video coding (DVC) is a video coding paradigm that allows for a low-complexity encoding process by exploiting the temporal redundancies in a video sequence at the decoder side. State-of-the-art DVC systems exhibit a structural coding delay since exploiting the temporal redundancies through motion-compensated interpolation requires the frames to be decoded out of order. To alleviate this problem, we propose a system based on motion-compensated extrapolation that allows for efficient low-delay video coding with low complexity at the encoder. The proposed extrapolation technique first estimates the motion field between the two most recently decoded frames using the Lucas–Kanade algorithm. The obtained motion field is then extrapolated to the current frame using an extrapolation grid. The proposed techniques are implemented into a novel architecture featuring hybrid block-frequency Wyner–Ziv coding as well as mode decision. Results show that having references from both temporal directions in interpolation provides superior rate-distortion performance over a single temporal direction in extrapolation, as expected. However, the proposed extrapolation method is particularly suitable for low-delay coding as it performs better than H.264/AVC intra, and it is even able to outperform the interpolation-based DVC codec from DISCOVER for several sequences.
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Christos Grecos, Adrian Munteanu 0001, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.4
2012 Distributed Video Coding With Feedback Channel Constraints
abstract
Many of the distributed video coding (DVC) systems described in the literature make use of a feedback channel from the decoder to the encoder to determine the rate. However, the number of requests through the feedback channel is often high, and as a result the overall delay of the system could be unacceptable in practical applications. As a solution, feedback-free DVC systems have been proposed, but the problem with these solutions is that they incorporate a difficult trade-off between encoder complexity and compression performance. Recognizing that a limited form of feedback may be supported in many video-streaming scenarios, in this paper we propose a method for constraining the number of feedback requests to a fixed maximum number of$N$requests for an entire Wyner-Ziv (WZ) frame. The proposed technique estimates the WZ rate at the decoder using information obtained from previously decoded WZ frames and defines the$N$requests by minimizing the expected rate overhead. Tests on eight sequences show that the rate penalty is less than 5% when only five requests are allowed per WZ frame (for a group of pictures of size four). Furthermore, due to improvements from previous work, the system is able to perform better than or similar to DISCOVER even when up to two requests per WZ frame are allowed. The practical usefulness of the proposed approach is studied by estimating end-to-end delay and encoder buffer requirements, indicating that DVC with constrained feedback can be an important solution in the context of video-streaming scenarios.
Jürgen Slowack, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
IEEE Trans. Circuits Syst. Video Technol.3
2012 Side-Information-Dependent Correlation Channel Estimation in Hash-Based Distributed Video Coding
abstract
In the context of low-cost video encoding, distributed video coding (DVC) has recently emerged as a potential candidate for uplink-oriented applications. This paper builds on a concept of correlation channel (CC) modeling, which expresses the correlation noise as being statistically dependent on the side information (SI). Compared with classical side-information-independent (SII) noise modeling adopted in current DVC solutions, it is theoretically proven that side-information-dependent (SID) modeling improves the Wyner-Ziv coding performance. Anchored in this finding, this paper proposes a novel algorithm for online estimation of the SID CC parameters based on already decoded information. The proposed algorithm enables bit-plane-by-bit-plane successive refinement of the channel estimation leading to progressively improved accuracy. Additionally, the proposed algorithm is included in a novel DVC architecture that employs a competitive hash-based motion estimation technique to generate high-quality SI at the decoder. Experimental results corroborate our theoretical gains and validate the accuracy of the channel estimation algorithm. The performance assessment of the proposed architecture shows remarkable and consistent coding gains over a germane group of state-of-the-art distributed and standard video codecs, even under strenuous conditions, i.e., large groups of pictures and highly irregular motion content.
Nikos Deligiannis, Joeri Barbarien, Adrian Munteanu 0001, Athanassios N. Skodras, Peter Schelkens
IEEE Trans. Image Process.1
2011 Distributed coding of endoscopic video
abstract
Triggered by the challenging prerequisites of wireless capsule endoscopic video technology, this paper presents a novel distributed video coding (DVC) scheme, which employs an original hash-based side-information creation method at the decoder. In contrast to existing DVC schemes, the proposed codec generates high quality side-information at the decoder, even under the strenuous motion conditions encountered in endoscopic video. Performance evaluation using broad endoscopic video material shows that the proposed approach brings notable and consistent compression gains over various state-of-the-art video codecs at the additional benefit of vastly reduced encoding complexity.
Nikos Deligiannis, Frederik Verbist, Joeri Barbarien, Jürgen Slowack, Rik Van de Walle, Peter Schelkens, Adrian Munteanu 0001
ICIP1
2011 Intra-WZ quantization mismatch in distributed video coding
abstract
During the past decade, Distributed Video Coding (DVC) has emerged as a new video coding paradigm, shifting the complexity from the encoder - to the decoder-side. This paper addresses a problem of current DVC architectures that has not been studied in the literature so far, that is, the mismatch between the intra and Wyner-Ziv (WZ) quantization processes. Due to this mismatch, WZ rate is spent even for spatial regions that are accurately approximated by the side-information. As a solution, this paper proposes side-information generation using selective unidirectional motion compensation from temporally adjacent WZ frames. Experimental results show that the proposed approach yields promising WZ rate gains of up to 7% relative to the conventional method.
Jürgen Slowack, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
ICIP5
2010 Bitplane intra coding with decoder-side mode decision in distributed video coding
abstract
While distributed video coding (DVC) has emerged as a new video coding paradigm, the compression performance of current systems is still low compared to conventional solutions such as H.264/AVC. While the latter uses many coding modes and an efficient mode decision strategy for choosing the best mode, in DVC, only a limited number of modes has been developed so far. Since encoder-side mode decision in DVC increases encoder's complexity, in this paper, we introduce decoder-side mode decision choosing between bitplane WZ coding and bitplane intra coding. This strategy proves to be efficient, delivering rate gains up to 22% over DISCOVER, without increasing the complexity at the encoder.
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
ICIP6
2010 Compensating for Motion Estimation Inaccuracies in DVC
Jürgen Slowack, Jozef Skorupa, Stefaan Mys, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
ICISP4
2010 Correlation modeling with decoder-side quantization distortion estimation for distributed video coding
abstract
Aiming for low-complexity encoding, distributed video coders still fail to achieve the performance of current industrial standards for video coding. One of most important problems in this area is the accurate modeling of the correlation between the predicted signal and the original video. In our previous work we showed that exploiting the quantization distortion can significantly improve the accuracy of a correlation estimator. In this paper we describe how the quantization distortion can be exploited purely at the decoder side without any performance penalty when compared to an encoder-aided system. As a result, the proposed correlation estimator delivers state-of-the-art modeling accuracy while neatly fitting the low-encoder-complexity characteristic of distributed video coding.
Jozef Skorupa, Jan De Cock, Jürgen Slowack, Stefaan Mys, Peter Lambert, Rik Van de Walle, Nikos Deligiannis, Adrian Munteanu 0001
PCS7
2010 Exploiting quantization and spatial correlation in virtual-noise modeling for distributed video coding
Jozef Skorupa, Jürgen Slowack, Stefaan Mys, Nikos Deligiannis, Jan De Cock, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Signal Process. Image Commun.4
2010 Rate-distortion driven decoder-side bitplane mode decision for distributed video coding
Jürgen Slowack, Stefaan Mys, Jozef Skorupa, Nikos Deligiannis, Peter Lambert, Adrian Munteanu 0001, Rik Van de Walle
Signal Process. Image Commun.4
2009 Modeling the Correlation Noise in Spatial Domain Distributed Video Coding
abstract
The paper thoroughly validates the proposed SID model on a broad dataset, showing significant accuracy improvements over SII models. Additionally, it is theoretically demonstrated that DVC systems that make SII assumptions suffer a performance penalty depending entirely on the correlation statistics of the video data.
Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Peter Schelkens, Jan Cornelis 0001
DCC1
2009 On the side-information dependency of the temporal correlation in Wyner-Ziv video coding
abstract
Current models in Wyner-Ziv video coding consider the temporal correlation noise to be side-information independent (SII). This paper goes beyond this assumption and proposes a novel model, of which the parameters are side-information dependent (SID). The proposed model is experimentally validated showing remarkable accuracy improvement over the conventional SII model. Moreover, a novel SID technique for the accurate estimation of the correlation channel in video is introduced. The proposed technique enables the design of a novel pixel-domain Wyner-Ziv video coding system operating without a feedback channel. Preliminary experimental results show that the proposed codec achieves superior performance compared to the state-of-the-art in pixel-domain Wyner-Ziv coding.
Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Jan Cornelis 0001, Peter Schelkens
ICASSP1
2009 Overlapped Block Motion Estimation and Probabilistic Compensation with Application in Distributed Video Coding
abstract
It has been recently demonstrated that, in distributed video coding (DVC) side-information dependent modeling of the correlation channel brings significant performance gains over side-information independent assumptions. In this letter, we present a novel technique enabling advanced side-information dependent estimation of the correlation channel at the decoder starting from a very coarse knowledge of it. The proposed technique triggers the design of a spatial-domain DVC codec which enables the suppression of the feedback channel. Experimental results show that the proposed codec achieves similar performance compared to the state-of-the-art in transform-domain Wyner-Ziv video coding while still operating at a fraction of its encoding complexity.
Nikos Deligiannis, Adrian Munteanu 0001, Tom Clerckx, Jan Cornelis 0001, Peter Schelkens
IEEE Signal Process. Lett.1