Danilo P. Mandic

dblp:41/6002 · DBLP profile ↗
← Back
288ranked-venue papers
23as first author
63since 2021 · last 2026
0000-0001-8432-3963ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 153 · 12 first-author · 30 since 2021Artificial intelligence and machine learning · 119 · 11 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 8Systems, architecture and hardware · 3Computer networks · 3Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
abstract
Developing open-set classification methods capable of classifying in-distribution (ID) data while detecting out-of-distribution (OOD) samples is essential for deploying graph neural networks (GNNs) in open-world scenarios. Existing methods typically treat all OOD samples as a single class, despite real-world applications—especially high-stake settings like fraud detection and medical diagnosis—demanding deeper insights into OOD samples, including their probable labels. This raises a critical question: Can OOD detection be extended to OOD classification without true label information? To answer this question, we introduce a Coarse-to-Fine open-set Classification (CFC) method that leverages large language models (LLMs) for text-attributed graphs. CFC consists of three key components: (1) A coarse classifier that utilizes LLM prompts for OOD detection and outlier label generation; (2) A GNN-based fine classifier trained with OOD samples from (1) for enhanced OOD detection and ID classification; and (3) Refined OOD classification achieved through LLM prompts and post-processed OOD labels. Unlike methods relying on synthetic or auxiliary OOD samples, CFC employs semantic OOD data-instances that are genuinely out-of-distribution based on their inherent meaning, thus improving interpretability and practical utility. CFC enhances OOD detection by 10% compared to state-of-the-art approaches on text-attributed graphs and in the text domain, while achieving up to 70% accuracy in OOD classification on graph datasets.
Xueqi Ma, Xingjun Ma, Sarah M. Erfani, Danilo P. Mandic, James Bailey 0001
AAAI4
2026 TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models
abstract
Parameter-Efficient Fine-Tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), have significantly reduced the number of trainable parameters needed in fine-tuning large language models (LLMs).The developments of LoRA-style adapters have considered two main directions: (1) enhancing model expressivity with high-rank adapters, and (2) aiming for further parameter reduction, as exemplified by vector-based methods.However, these approaches come with a trade-off, as achieving the expressivity of high-rank weight updates typically comes at the cost of sacrificing the extreme parameter efficiency offered by vector-based techniques.To address this issue, we propose a vector-based random Tensor network for high-Rank Adaptation (TeRA), a novel PEFT method that achieves high-rank weight updates while retaining the parameter efficiency of vector-based PEFT adapters.This is achieved by parametrizing the tensorized weight update matrix as a Tucker-like tensor network (TN), whereby large randomly initialized factors are frozen and shared across layers, while only small layer-specific scaling vectors, corresponding to diagonal entries of factor matrices, are trained.Comprehensive experiments demonstrate that TeRA matches or even outperforms existing high-rank adapters, while requiring as few trainable parameters as vectorbased methods.Theoretical analysis and ablation studies validate the effectiveness of the proposed TeRA method.The code is available at https://github.com/guyuxuan9/TeRA.
Yuxuan Gu 0005, Wuyang Zhou, Giorgos Iacovides, Danilo P. Mandic
ACL (1)4
2026 Quaternion-based CNN for heart rate prediction from PPG
Youngshin Kang, Yusang Nam, Hyuntae Lee, Danilo P. Mandic
Neural Networks5
2026 A novel backpropagation algorithm based on negated kurtosis loss for training shallow, convolutional, and deep neural networks
Engin Cemal Menguc, Alper Emlek, Danilo P. Mandic
Neural Networks3
2025 Kernel-Based Anomaly Detection Using Generalized Hyperbolic Processes
abstract
We present a novel approach to anomaly detection by integrating Generalized Hyperbolic (GH) processes into kernel-based methods. The GH distribution, known for its flexibility in modeling skewness, heavy tails, and kurtosis, helps to capture complex patterns in data that deviate from Gaussian assumptions. We propose a GH-based kernel function and utilize it within Kernel Density Estimation (KDE) and One-Class Support Vector Machines (OCSVM) to develop anomaly detection frameworks. Theoretical results confirmed the positive semi-definiteness and consistency of the GH-based kernel, ensuring its suitability for machine learning applications. Empirical evaluation on synthetic and real-world datasets showed that our method improves detection performance in scenarios involving heavy-tailed and asymmetric or imbalanced distributions. https://github.com/paulinebourigault/GHKernelAnomalyDetect.
Pauline Bourigault, Danilo P. Mandic
ICASSP2
2025 RespDiff: An End-to-End Multi-scale RNN Diffusion Model for Respiratory Waveform Estimation from PPG Signals
abstract
Respiratory rate (RR) is a critical health indicator often monitored under inconvenient scenarios, limiting its practicality for continuous monitoring. Photoplethysmography (PPG) sensors, increasingly integrated into wearable devices, offer a chance to continuously estimate RR in a portable manner. In this paper, we propose RespDiff, an end-to-end multi-scale RNN diffusion model for respiratory waveform estimation from PPG signals. RespDiff does not require hand-crafted features or the exclusion of low-quality signal segments, making it suitable for real-world scenarios. The model employs multi-scale encoders, to extract features at different resolutions, and a bidirectional RNN to process PPG signals and generate respiratory waveform. Additionally, a spectral loss term is introduced to optimize the model further. Experiments conducted on the BIDMC dataset demonstrate that RespDiff outperforms notable previous works, achieving a mean absolute error (MAE) of 1.18 bpm for RR estimation while others range from 1.66 to 2.15 bpm, showing its potential for robust and accurate respiratory monitoring in real-world applications.
Yuyang Miao, Zehua Chen 0005, Danilo P. Mandic
ICASSP4
2025 Panorama: An enabling technology for Hearables
abstract
The emergence of in-ear wearable sensors, so called Hearables, offers a convenient and cost-effective tool for long-term health monitoring. Traditionally, extracting frequency domain features from these signals relies on Power Spectral Density (PSD), which is highly susceptible to the effects of stochastic noise. This poses a formidable challenge for the processing of in-ear biosignals, which often have low signal-to-noise ratios (SNR). To address this problem, this work employs our recently proposed method for spectral analysis, termed "Panorama", defined as the Fourier transform of the autoconvolution. Experiments on in-ear EEG signals with extremely low SNR and in-ear sleep classification tasks show that, compared with the widely-used PSD, the Panorama spectrum offers better separability of features. We also show that a time series can be reconstructed from its autoconvolution, the inverse Fourier transform of Panorama, offering an additional analysis tool for time-domain applications.
Qiyu Rao, Zdenka Babic, Scott C. Douglas, Danilo P. Mandic
ICASSP4
2025 Relational Conformal Prediction for Correlated Time Series
abstract
We address the problem of uncertainty quantification in time series forecasting by exploiting observations at correlated sequences. Relational deep learning methods leveraging graph representations are among the most effective tools for obtaining point estimates from spatiotemporal data and correlated time series. However, the problem of exploiting relational structures to estimate the uncertainty of such predictions has been largely overlooked in the same context. To this end, we propose a novel distribution-free approach based on the conformal prediction framework and quantile regression. Despite the recent applications of conformal prediction to sequential data, existing methods operate independently on each target time series and do not account for relationships among them when constructing the prediction interval. We fill this void by introducing a novel conformal prediction method based on graph deep learning operators. Our approach, named Conformal Relational Prediction (CoRel), does not require the relational structure (graph) to be known a priori and can be applied on top of any pre-trained predictor. Additionally, CoRel includes an adaptive component to handle non-exchangeable data and changes in the input time series. Our approach provides accurate coverage and achieves state-of-the-art uncertainty quantification in relevant benchmarks.
Andrea Cini, Alexander Jenkins, Danilo P. Mandic, Cesare Alippi, Filippo Maria Bianchi
ICML3
2025 TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs
abstract
The reasoning abilities of Large Language Models (LLMs) can be improved by structurally denoising their weights, yet existing techniques primarily focus on denoising the feed-forward network (FFN) of the transformer block, and can not efficiently utilise the Multi-head Attention (MHA) block, which is the core of transformer architectures. To address this issue, we propose a novel intuitive framework that, at its very core, performs MHA compression through a multi-head tensorisation process and the Tucker decomposition. This enables both higher-dimensional structured denoising and compression of the MHA weights, by enforcing a shared higher-dimensional subspace across the weights of the multiple attention heads. We demonstrate that this approach consistently enhances the reasoning capabilities of LLMs across multiple benchmark datasets, and for both encoder-only and decoder-only architectures, while achieving compression rates of up to ~ 250 times in the MHA weights, all without requiring any additional data, training, or fine-tuning. Furthermore, we show that the proposed method can be seamlessly combined with existing FFN-only-based denoising techniques to achieve further improvements in LLM reasoning performance.1
Yuxuan Gu 0005, Wuyang Zhou, Giorgos Iacovides, Danilo P. Mandic
IJCNN4
2025 Augmentation of EEG and ECG Time Series for Machine Learning Applications: Integrating Changepoint Detection into the iAAFT Surrogates
abstract
The performance of deep learning methods is dependent on the quality and quantity of the available training data, in particular for physiological time series data which is scarce and noisy. Thus, data augmentation methods are used to artificially increase the size of datasets. However, the time-evolving statistical properties of nonstationary signals prevent the use of standard data augmentation techniques. To this end, we introduce a novel method for augmenting nonstationary time series. The method combines offline changepoint detection with the iterative amplitude-adjusted Fourier transform (iAAFT), thereby ensuring that the time-frequency properties of the original signal are pre-served during augmentation. The proposed method is validated by comparing the performance of i) a deep learning seizure detection algorithm on both the original and augmented versions of the CHB-MIT and Siena scalp electroencephalography (EEG) datasets, ii) a feature-based machine learning seizure detection algorithm on both the original and augmented versions of the CHB-MIT and Siena scalp EEG datasets, and iii) a deep learning atrial fibrillation (AF) detection algorithm on the original and augmented versions of the Computing in Cardiology Challenge 2017 dataset. For deep learning classification on the CHB-MIT and Siena datasets, respectively, the proposed method increased accuracy by 2.2% and 4.6%, precision by 2.3% and 3.6%, recall by 2.8% and 2.3%, and F1 by 4.2% and 4.2%. For feature-based classification using the CHB-MIT and Siena datasets respectively, accuracy increased by 2.2% and 0.2%, precision by 4.3% and 0.2%, recall by 5.2% and 0.1%, and F1 by 5.4% and 0.3%. For AF classification, the accuracy increased by 0.5%, precision by 6%, recall by 0.1%, and F1 by 2.8%.
Nina Moutonnet, Gregory Scott, Danilo P. Mandic
IJCNN3
2025 Online censoring-based learning algorithms for fully complex-valued neural networks
Engin Cemal Menguc, Danilo P. Mandic
Neurocomputing2
2025 Tensor ring rank determination using odd-dimensional unfolding
abstract
While tensor ring (TR) decomposition methods have been extensively studied, the determination of TR-ranks remains a challenging problem, with existing methods being typically sensitive to the determination of the starting rank (i.e., the first rank to be optimized). Moreover, current methods often fail to adaptively determine TR-ranks in the presence of noisy and incomplete data, and exhibit computational inefficiencies when handling high-dimensional data. To address these issues, we propose an odd-dimensional unfolding method for the effective determination of TR-ranks. This is achieved by leveraging the symmetry of the TR model and the bound rank relationship in TR decomposition. In addition, we employ the singular value thresholding algorithm to facilitate the adaptive determination of TR-ranks and use randomized sketching techniques to enhance the efficiency and scalability of the method. Extensive experimental results in rank identification, data denoising, and completion demonstrate the potential of our method for a broad range of applications.
Yichun Qiu, Guoxu Zhou, Chao Li 0013, Danilo P. Mandic, Qibin Zhao
Neural Networks4
2025 Optimizing beamforming in quaternion signal processing using projected gradient descent algorithm
Qiankun Diao, Dongpo Xu, Shuning Sun, Danilo P. Mandic
Signal Process.4
2025 A class of widely linear quaternion blind equalisation algorithms
Min Xiang, Sayed Pouria Talebi, Danilo P. Mandic
Signal Process.5
2024 Multi-Modal Information Bottleneck Attribution with Cross-Attention Guidance
Pauline Bourigault, Emmanuelle Bourigault, Danilo P. Mandic
BMVC3
2024 Segmented Error Minimisation (Semi) for Robust Training of Deep Learning Models with Non-Linear Shifts in Reference Data
abstract
Time series regression models are typically trained using the mean squared error (MSE) and thus rely critically on time-aligned reference data. However, the MSE loss is often inadequate when processing real-world data, such as physiological signals, as misalignment between two signals can cause a large change in MSE, thus severely inhibiting convergence. This is regularly compounded by time varying drifts between different modalities and such as photoplethysmography, electrocardiography, blood pressure and respiration. Indeed, these exhibit fluctuating time delays even when taken from the same individual across different body positions and at different times of the day. To this end, we introduce the concept of segmented error minimisation (SEMI), a new loss function which accounts for differing time delays among the variables. The SEMI is examined both through simulations, and via a denoising convolutional autoencoder with synthetic data. It is then finally verified in the real-world application of the denoising of wearable photoplethysmography with a reference signal.
Harry J. Davies, Yuyang Miao, Amir Nassibi, Morteza Khaleghimeybodi, Danilo P. Mandic
ICASSP5
2024 Graph-based Time Series Clustering for End-to-End Hierarchical Forecasting
abstract
Relationships among time series can be exploited as inductive biases in learning effective forecasting models. In hierarchical time series, relationships among subsets of sequences induce hard constraints (hierarchical inductive biases) on the predicted values. In this paper, we propose a graph-based methodology to unify relational and hierarchical inductive biases in the context of deep learning for time series forecasting. In particular, we model both types of relationships as dependencies in a pyramidal graph structure, with each pyramidal layer corresponding to a level of the hierarchy. By exploiting modern - trainable - graph pooling operators we show that the hierarchical structure, if not available as a prior, can be learned directly from data, thus obtaining cluster assignments aligned with the forecasting objective. A differentiable reconciliation stage is incorporated into the processing architecture, allowing hierarchical constraints to act both as an architectural bias as well as a regularization element for predictions. Simulation results on representative datasets show that the proposed method compares favorably against the state of the art.
Andrea Cini, Danilo P. Mandic, Cesare Alippi
ICML2
2024 Quaternion Recurrent Neural Network with Real-Time Recurrent Learning and Maximum Correntropy Criterion
abstract
We develop a robust quaternion recurrent neural network (QRNN) for real-time processing of 3D and 4D data with outliers. This is achieved by combining the real-time recurrent learning (RTRL) algorithm and the maximum correntropy criterion (MCC) as a loss function. While both the mean square error and maximum correntropy criterion are viable cost functions, it is shown that the non-quadratic maximum correntropy loss function is less sensitive to outliers, making it suitable for applications with multidimensional noisy or uncertain data. Both algorithms are derived based on the novel generalised HR (GHR) calculus, which allows for the differentiation of real functions of quaternion variables and offers the product and chain rules, thus enabling elegant and compact derivations. Simulation results in the context of motion prediction of chest internal markers for lung cancer radiotherapy, which includes regular and irregular breathing sequences, support the analysis.
Pauline Bourigault, Dongpo Xu, Danilo P. Mandic
IJCNN3
2024 GPINN: Physics-Informed Neural Network with Graph Embedding
abstract
Solving partial differential equations (PDEs) is of great importance in numerous fields including physics, engineering, finance, and scientific computing. Physics-Informed Neural Networks (PINNs) have gained interest due to their ability to solve PDEs in the strong form, which differs from the traditional methods like the Finite Element Method (FEM) which exhibits the weak form. However, the PINN methods normally operate in the Euclidean space only and lack spatial contextual awareness. This deficiency leads to distance inequality, whereby the Euclidean distance might not align with the actual physical distance. Such misalignment may lead the model to produce high errors with no physical meanings. In this study, we propose the Physics-Informed Neural Network with Graph Embedding (GPINN), which utilises the eigenvectors of graph Laplacian to transform the input space from a pure Euclidean space to a joint Euclidean and topological (graph-based) space and make the model spatial-context aware. Two case studies are conducted to model heat propagation and linear elasticity and show the enhancement from PINN to GPINN.
Yuyang Miao, Danilo P. Mandic
IJCNN3
2024 Widely Linear Matched Filter: A Lynchpin towards the Interpretability of Complex-valued CNNs
abstract
A recent study on the interpretability of real-valued convolutional neural networks (CNNs) [1] has revealed a direct and physically meaningful link with the task of finding features in data through matched filters. However, applying this paradigm to illuminate the interpretability of complex-valued CNNs meets a formidable obstacle: the extension of matched filtering to a general class of noncircular complex-valued data, referred to here as the widely linear matched filter (WLMF), has been only implicit in the literature. To this end, to establish the interpretability of the operation of complex-valued CNNs, we introduce a general WLMF paradigm, provide its solution and undertake analysis of its performance. For rigor, our WLMF solution is derived without imposing any assumption on the probability density of noise. The theoretical advantages of the WLMF over its standard strictly linear counterpart (SLMF) are provided in terms of their output signal-to-noise-ratios (SNRs), with WLMF consistently exhibiting enhanced SNR. Moreover, the lower bound on the SNR gain of WLMF is derived, together with condition to attain this bound. This serves to revisit the convolution-activation-pooling chain in complex-valued CNNs through the lens of matched filtering, which reveals the potential of WLMFs to provide physical interpretability and enhance explainability of general complex-valued CNNs. Simulations demonstrate the agreement between the theoretical and numerical results.
Qingchen Wang, Zhe Li 0007, Zdenka Babic, Ljubisa Stankovic, Danilo P. Mandic
IJCNN6
2024 UAdam: Unified Adam-Type Algorithmic Framework for Nonconvex Optimization
abstract
Adam-type algorithms have become a preferred choice for optimization in the deep learning setting; however, despite their success, their convergence is still not well understood. To this end, we introduce a unified framework for Adam-type algorithms, termed UAdam. It is equipped with a general form of the second-order moment, which makes it possible to include Adam and its existing and future variants as special cases, such as NAdam, AMSGrad, AdaBound, AdaFom, and Adan. The approach is supported by a rigorous convergence analysis of UAdam in the general nonconvex stochastic setting, showing that UAdam converges to the neighborhood of stationary points with a rate of O(1/T). Furthermore, the size of the neighborhood decreases as the parameter β1 increases. Importantly, our analysis only requires the first-order momentum factor to be close enough to 1, without any restrictions on the second-order momentum factor. Theoretical results also reveal the convergence conditions of vanilla Adam, together with the selection of appropriate hyperparameters. This provides a theoretical guarantee for the analysis, applications, and further developments of the whole general class of Adam-type algorithms. Finally, several numerical experiments are provided to support our theoretical findings.
Yiming Jiang 0013, Dongpo Xu, Danilo P. Mandic
Neural Comput.4
2024 Finding core labels for maximizing generalization of graph neural networks
Sichao Fu, Xueqi Ma, Yibing Zhan, Fanyu You, Qinmu Peng, Tongliang Liu, James Bailey 0001, Danilo P. Mandic
Neural Networks8
2024 Price's Theorem for Quaternion Variables
abstract
Price's theorem in statistical signal processing relates the expectation of a nonlinear function of normally distributed random variables to their covariances. However, such a key theorem is yet to be established and explored in quaternion statistics. To this end, we introduce Price's theorem for quaternion variables using the generalized Hamilton-real (GHR) calculus. This is achieved by first employing the chain rule of GHR calculus to derive two crucial quaternion matrix derivatives of functions with respect to the product of quaternion matrix variables. Next, we leverage quaternion second-order statistics to establish the relationship between the derivative of a function with respect to the augmented quaternion covariance matrix and its real counterpart. Based on the above results and Price's theorem in real signal processing, we finally propose a novel formulation of Price's theorem for quaternion random variables. This finding not only enriches the theory of quaternion statistical signals processing but also extends its applicability.
Qiankun Diao, Dongpo Xu, Shuning Sun, Danilo P. Mandic
IEEE Signal Process. Lett.4
2024 Introduction to the Special Issue on AI-Generated Content for Multimedia
abstract
Our world is becoming rapidly dependent on data of increasing complexity, diversity, and volume which calls for robust and powerful tools to process such big data. Probabilistic generative models fulfill this goal by learning latent characteristic data relations, especially for the recent emergence of large-scale deep generative models that are able to create realistic content, namely, artificial intelligence-generated content (AIGC). The applications of AIGC span across various domains, and witness rich potential in multimedia content creation, including dialog generation, text-to-speech conversion, image/video generation, and cross-modal content generation.
Shengxi Li, Xuelong Li 0001, Leonardo Chiariglione, Jiebo Luo 0001, Wenwu Wang 0001, Zhengyuan Yang, Danilo P. Mandic, Hamido Fujita
IEEE Trans. Circuits Syst. Video Technol.7
2024 Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural Networks
abstract
Graph neural networks (GNNs) tend to suffer from high computation costs due to the exponentially increasing scale of graph data and a large number of model parameters, which restricts their utility in practical applications. To this end, some recent works focus on sparsifying GNNs (including graph structures and model parameters) with the lottery ticket hypothesis (LTH) to reduce inference costs while maintaining performance levels. However, the LTH-based methods suffer from two major drawbacks: 1) they require exhaustive and iterative training of dense models, resulting in an extremely large training computation cost, and 2) they only trim graph structures and model parameters but ignore the node feature dimension, where vast redundancy exists. To overcome the above limitations, we propose a comprehensive graph gradual pruning framework termed CGP. This is achieved by designing a during-training graph pruning paradigm to dynamically prune GNNs within one training process. Unlike LTH-based methods, the proposed CGP approach requires no retraining, which significantly reduces the computation costs. Furthermore, we design a cosparsifying strategy to comprehensively trim all the three core elements of GNNs: graph structures, node features, and model parameters. Next, to refine the pruning operation, we introduce a regrowth process into our CGP framework, to reestablish the pruned but important connections. The proposed CGP is evaluated over a node classification task across six GNN architectures, including shallow models [graph convolutional network (GCN) and graph attention network (GAT)], shallow-but-deep-propagation models [simple graph convolution (SGC) and approximate personalized propagation of neural predictions (APPNP)], and deep models [GCN via initial residual and identity mapping (GCNII) and residual GCN (ResGCN)], on a total of 14 real-world graph datasets, including large-scale graph datasets from the challenging Open Graph Benchmark (OGB). Experiments reveal that the proposed strategy greatly improves both training and inference efficiency while matching or even exceeding the accuracy of the existing methods.
Chuang Liu 0008, Xueqi Ma, Yibing Zhan, Liang Ding 0006, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.8
2024 Matched Filtering on Directed Graphs
abstract
The matched filter is a crucial concept in both signal analysis and convolutional neural networks (CNNs). Previous work has addressed graph matched filtering principles for undirected graphs. This article expands upon the existing literature, by exploring matched filtering principles for signals on directed graphs. In such cases, the adjacency matrix is asymmetric, and commonly results in nonorthogonal eigenvectors. The presented concept is supported by a detailed analysis and numerical examples.
Isidora Stankovic, Milos Brajovic, Cornel Ioana, Milos Dakovic, Danilo P. Mandic, Ljubisa Stankovic
IEEE Trans. Syst. Man Cybern. Syst.5
2023 Hierarchical Graph Learning for Stock Market Prediction Via a Domain-Aware Graph Pooling Operator
abstract
The utility of Graph Neural Networks (GNN) for the paradigm of forecasting short-term stock price movements is investigated. In particular, a finance-specific graph pooling operation, referred to as StockPool, is introduced to efficiently coarsen the stock graph. This is achieved by employing domain knowledge to cluster stocks, depending on some task-specific characteristics (e.g. industries, sub-industries, etc.). Unlike fully end-to-end learnable graph pooling strategies (e.g. differentiable pooling, MinCUT pooling, etc.), such a deterministic pooling operator is considerably more computationally efficient and thus scalable to larger stock graphs. Experimentations on the S&P500 stock index demonstrate that the StockPool operator outperforms existing graph pooling strategies on the prediction of price movements. Finally, different graph pooling methods are utilized to create a set of highly uncorrelated GNN models; these are used to construct a graph ensemble model with an improved performance.
Arie N. Arya, Yao Lei Xu, Ljubisa Stankovic, Danilo P. Mandic
ICASSP4
2023 Tensor Completion for Efficient and Accurate Hyperparameter Optimisation in Large-Scale Statistical Learning
abstract
Hyperparameter optimisation is a prerequisite for state-of-the- art performance in machine learning, with current strategies including Bayesian optimisation, hyperband, and evolutionary methods. While such methods have been shown to improve performance, none of these is designed to explicitly take advantage of the underlying data structure. To this end, we introduce a completely different approach for hyperparameter optimisation, based on low-rank tensor completion. This is achieved by first forming a multi-dimensional tensor which comprises performance scores for different combinations of hyperparameters. Based on the realistic underlying assumption that the so-formed tensor has a low-rank structure, reliable estimates of the unobserved validation scores of combinations of hyper- parameters are next obtained through tensor completion, from only a fraction of the known elements in the tensor. Through extensive experimentation on various datasets and learning models, the proposed method is shown to exhibit competitive or superior performance to state-of-the-art hyperparameter optimisation strategies. Distinctive advantages of the proposed method include its ability to simultaneously handle any hyper- parameter type (kind of optimiser, number of neurons, number of layer, etc.), its relative simplicity compared to competing methods, as well as the ability to suggest multiple optimal combinations of hyperparameters.
Aaman Rebello, Kriton Konstantinidis, Yao Lei Xu, Danilo P. Mandic
ICASSP4
2023 Relating EEG Recordings to Speech Using Envelope Tracking and The Speech-FFR
abstract
During speech perception, a listener’s electroencephalogram (EEG) reflects acoustic-level processing as well as higher-level cognitive factors such as speech comprehension and attention. However, decoding speech from EEG recordings is challenging due to the low signal-to-noise ratios of EEG signals. We report on an approach developed for the ICASSP 2023 ‘Auditory EEG Decoding’ Signal Processing Grand Challenge. A simple ensembling method is shown to considerably improve upon the baseline decoder performance. Even higher classification rates are achieved by jointly decoding the speech-evoked frequency-following response and responses to the temporal envelope of speech, as well as by fine-tuning the decoders to individual subjects. Our results could have applications in the diagnosis of hearing disorders or in cognitively steered hearing aids.
Mike Thornton, Danilo P. Mandic, Tobias Reichenbach
ICASSP2
2023 ClassA Entropy for the Analysis of Structural Complexity of Physiological Signals
abstract
Despite the recent theoretical boom in Sample Entropy based algorithms for the analysis of physiological and pathological systems, the major issue which prevents their more widespread use remains that of large computational load, particularly in the studies of quantification of structural richness in data. This issue becomes even more prohibitive when it comes to large data sizes and real-time processing. To this end, a new Classification Angle (ClassA) Entropy is introduced for the quantification of structural complexity of real world signals, based on an improved Second Order Difference Plot and Shannon Entropy. In comparison with existing nonlinear techniques, including Real Sum Angle index, Phase Entropy, Gridded Distribution Entropy and Cosine Similarity Entropy, the proposed method offers the advantages of a minimal number of tunable parameters, lower requirement of data size, wider application field and large relaxation of computational load. Simulations on real world physiological data support the approach.
Hongjian Xiao, Ling Li 0010, Danilo P. Mandic
ICASSP3
2023 AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
abstract
Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn continuous audio representations from contrastive language-audio pretraining (CLAP) embeddings. The pretrained CLAP models enable us to train LDMs with audio embeddings while providing text embeddings as the condition during sampling. By learning the latent representations of audio signals without modelling the cross-modal relationship, AudioLDM improves both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance compared to other open-sourced systems, measured by both objective and subjective metrics. AudioLDM is also the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at https://audioldm.github.io.
Haohe Liu, Zehua Chen 0005, Xinhao Mei, Xubo Liu 0001, Danilo P. Mandic, Wenwu Wang 0001, Mark D. Plumbley
ICML6
2023 Last-iterate convergence analysis of stochastic momentum methods for neural networks
Dongpo Xu, Yinghua Lu, Danilo P. Mandic
Neurocomputing5
2023 Graph-Regularized Tensor Regression: A Domain-Aware Framework for Interpretable Modeling of Multiway Data on Graphs
abstract
Modern data analytics applications are increasingly characterized by exceedingly large and multidimensional data sources. This represents a challenge for traditional machine learning models, as the number of model parameters needed to process such data grows exponentially with the data dimensions, an effect known as the curse of dimensionality. Recently, tensor decomposition (TD) techniques have shown promising results in reducing the computational costs associated with large-dimensional models while achieving comparable performance. However, such tensor models are often unable to incorporate the underlying domain knowledge when compressing high-dimensional models. To this end, we introduce a novel graph-regularized tensor regression (GRTR) framework, whereby domain knowledge about intramodal relations is incorporated into the model in the form of a graph Laplacian matrix. This is then used as a regularization tool to promote a physically meaningful structure within the model parameters. By virtue of tensor algebra, the proposed framework is shown to be fully interpretable, both coefficient-wise and dimension-wise. The GRTR model is validated in a multiway regression setting and compared against competing models and is shown to achieve improved performance at reduced computational costs. Detailed visualizations are provided to help readers gain an intuitive understanding of the employed tensor operations.
Yao Lei Xu, Kriton Konstantinidis, Danilo P. Mandic
Neural Comput.3
2023 A class of doubly stochastic shift operators for random graph signals and their boundedness
Bruno Scalzo Dees, Ljubisa Stankovic, Milos Dakovic, Anthony G. Constantinides, Danilo P. Mandic
Neural Networks5
2023 Reciprocal GAN Through Characteristic Functions (RCF-GAN)
abstract
The integral probability metric (IPM) equips generative adversarial nets (GANs) with the necessary theoretical support for comparing statistical moments in an embedded domain of the critic, while stabilising their training and mitigating the mode collapse issues. For enhanced intuition and physical insight, we introduce a generalisation of IPM-GANs which operates by directly comparing probability distributions rather than their moments. This is achieved through characteristic functions (CFs), a powerful tool that uniquely comprises all information about any general distribution. For rigour, we first theoretically prove the ability of the CF loss to compare probability distributions, and proceed to establish the physical meaning of the phase and amplitude of CFs. An optimal sampling strategy is then developed to calculate the CFs, and an equivalence between the embedded and data domains is proved under the reciprocal theory. This makes it possible to seamlessly combine IPM-GAN with an auto-encoder structure by an advanced anchor architecture, which adversarially learns a semantic low-dimensional manifold for both generation and reconstruction. This efficient reciprocal CF GAN (RCF-GAN) structure, uses only two modules and a simple training strategy to achieve the state-of-the-art bi-directional generation. Experiments demonstrate the superior performance of RCF-GAN on both regular (images) and irregular (graph) domains.
Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Performance analysis of the augmented complex-valued least mean kurtosis algorithm
Jingen Ni, Zhe Li 0007, Engin Cemal Menguc, Jie Chen 0022, Danilo P. Mandic
Signal Process.6
2023 A Class of Online Censoring Based Quaternion-Valued Least Mean Square Algorithms
abstract
Streaming Big Data applications require the means to efficiently utilize large-scale data in an online manner. This issue becomes even more pressing when data are also multidimensional, as is the case with quaternion data streams. To this end, we first introduce the online censoring (OC) based quaternion least mean square (OC-QLMS) and OC-augmented QLMS (OC-AQLMS) algorithms, which censor less informative data in order to reduce computational complexity without severely affecting performance. Next, to censor both the outlier and noninformative data, we also propose the robust OC-QLMS (ROC-QLMS) and ROC-AQLMS. Fixed and adaptive threshold rules are introduced into the proposed OC algorithms to efficiently implement the desired censoring probability in the quaternion domain. The fundamental convergence analysis on the step size for all the proposed algorithms is also presented and the superior properties of the proposed algorithms are demonstrated in system identification scenarios.
Engin Cemal Menguc, Nurettin Acir, Danilo P. Mandic
IEEE Signal Process. Lett.3
2023 Von Mises-Fisher Elliptical Distribution
abstract
Modern probabilistic learning systems mainly assume symmetric distributions, however, real-world data typically obey skewed distributions and are thus not adequately modeled through symmetric distributions. To address this issue, a generalization of symmetric distributions called elliptical distributions are increasingly used, together with further improvements based on skewed elliptical distributions. However, existing approaches are either hard to estimate or have complicated and abstract representations. To this end, we propose a novel approach based on the von-Mises-Fisher (vMF) distribution to obtain an explicit and simple probability representation of skewed elliptical distributions. The analysis shows that this not only allows us to design and implement nonsymmetric learning systems but also provides a physically meaningful and intuitive way of generalizing skewed distributions. For rigor, the proposed framework is proven to share important and desirable properties with its symmetric counterpart. The proposed vMF distribution is demonstrated to be easy to generate and stable to estimate, both theoretically and through examples.
Shengxi Li, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.2
2023 Convolutional Neural Networks Demystified: A Matched Filtering Perspective-Based Tutorial
abstract
Deep neural networks (DNNs) and especially convolutional neural networks (CNNs) have revolutionized the way we approach the analysis of large quantities of data. However, the largely ad hoc fashion of their development, albeit one reason for their rapid success, has also brought to light the intrinsic limitations of CNNs—in particular, those related to their black box nature. In addition, the ability to “explain” both the way such systems behave and the results they produce is increasingly becoming an imperative in many practical applications. Therefore, it would be particularly useful to establish physically meaningful mechanisms underpinning the operation of CNNs, thus helping to resolve the issue of interpretability of the processing steps and explain their input-output relationship. To this end, we revisit the operation of CNNs from first principles and show that their very backbone—the convolution operation—represents a matched filter which examines the input for the presence of characteristic patterns in data. Our treatment is based on temporal signals, naturally generated by physical sensors, which admit rigorous analysis through systems science. This serves as a vehicle for a unifying account on the overall functionality of CNNs, whereby both the convolution-activation-pooling chain and learning strategies are shown to admit a compact and elegant interpretation under the umbrella of matched filtering. In addition to helping reveal the physical principles underpinning CNNs and providing an intuitive understanding of their operation, the treatment of CNNs from a matched filtering perspective is also shown to offer a platform to support further developments in this area.
Ljubisa Stankovic, Danilo P. Mandic
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Dynamic Portfolio Cuts: A Spectral Approach to Graph-Theoretic Diversification
abstract
Stock market returns are typically analyzed using standard regression models yet they reside on irregular domains, a natural scenario for graph signal processing. This motivates us to consider a market graph as an intuitive way to represent the relationships between financial assets. Traditional methods for estimating asset-return covariance operate under the assumption of statistical time-invariance, and are thus unable to appropriately infer the underlying structure of the market graph. To this end, this work introduces a class of graph spectral estimators which cater for the nonstationarity inherent to asset price movements, as a basis to represent the time-varying interactions between assets through a dynamic spectral market graph. Such an account of the time-varying nature of the asset-return covariance allows us to introduce the notion of dynamic spectral portfolio cuts, whereby the graph is partitioned into time-evolving clusters, thus allowing for robust and online asset allocation. The advantages of the proposed framework over traditional methods are demonstrated through numerical case studies using real-world price data.
Alvaro Arroyo, Bruno Scalzo Dees, Ljubisa Stankovic, Danilo P. Mandic
ICASSP4
2022 Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
abstract
Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to optimize the choice of inference schedule over a few iterations to speed up inference. However, this results in reduced generation quality, mainly because the inference process is optimized separately, without jointly optimizing with the training process. In this paper, we propose InferGrad, a diffusion model for vocoder that incorporates inference process into training, to reduce the inference iterations while maintaining high generation quality. More specifically, during training, we generate data from random noise through a reverse process under inference schedules with a few iterations, and impose a loss to minimize the gap between the generated and ground-truth data samples. Then, unlike existing approaches, the training of InferGrad considers the inference process. The advantages of InferGrad are demonstrated through experiments on the LJSpeech dataset showing that InferGrad achieves better voice quality than the baseline WaveGrad under same conditions while maintaining the same voice quality as the baseline but with 3x speedup (2 iterations for InferGrad vs 6 iterations for WaveGrad).
Zehua Chen 0005, Xu Tan 0003, Shifeng Pan, Danilo P. Mandic, Lei He 0005, Sheng Zhao 0002
ICASSP5
2022 Variational Bayesian Tensor Networks with Structured Posteriors
abstract
Tensor network (TN) methods have proven their considerable potential in deterministic regression and classification related paradigms, but remain underexplored in probabilistic settings. To this end, we introduce a variational inference framework for supervised learning in the context of TNs, referred to as the Bayesian Tensor Network (BTN). This is achieved by making use of the multi-linear nature of tensor networks which allows us to construct a structured variational model which scales linearly with data dimensionality. The so imposed low rank structure on the tensor mean and Kronecker separability of the local covariances makes it possible to efficiently induce weight dependencies in the posterior distribution. This is shown to enhance model expressiveness at a drastically lower parameter complexity compared to the standard mean-field approach. A comprehensive validation of the proposed framework demonstrates the competitiveness of BTNs against existing structured Bayesian neural network approaches, while exhibiting enhanced interpretability, computational efficiency, and ability to yield credibility intervals.
Kriton Konstantinidis, Yao Lei Xu, Qibin Zhao, Danilo P. Mandic
ICASSP4
2022 Multivariate Multiscale Cosine Similarity Entropy
abstract
The rapid development in sensor technology has made it convenient to acquire data from multi-channel systems but has also high-lighted the need for the analysis of nonlinear dynamical properties at a higher level - the so-called structural complexity. Traditional single-scale entropy measures, such as the amplitude based Sample Entropy (SampEn), are designed to give a quantification of irregularity and randomness. Its enhanced versions, Multiscale Sample Entropy (MSampEn) and Multivariate Multiscale Sample Entropy (MMSE), are capable of detecting the structure within a signal at high scales and for multivariate data, however, the scaling process comes at a cost of the reduction of the number of sample points that results in reduced stability and limitations regarding the selection of the embedding dimension. In addition, the analyses of structure on the basis of MSampEn and MMSE require relatively high scales, yet without prior-knowledge of the scale degree. To this end, we propose a new multivariate entropy method based on the recently introduced Cosine Similarity Entropy (CSE). The proposed Multivariate Multiscale Cosine Similarity Entropy (MMCSE) is based on angular distance which makes it possible to assess long-term correlation within a system at both a low and large scales, and thus assess the true structural complexity in a more physically meaningful way. Both synthetic and real world signals are utilized to examine the performance of the proposed approach, with the resulting simulations supporting the approach.
Hongjian Xiao, Theerasak Chanwimalueang, Danilo P. Mandic
ICASSP3
2022 Low-Complexity Attention Modelling via Graph Tensor Networks
abstract
The attention mechanism is at the core of modern Natural Language Processing (NLP) models, owing to its ability to focus on the most contextually relevant part of a sequence. However, current attention models rely on "flat-view" matrix methods to process tokens embedded in vector spaces; this results in exceedingly high parameter complexity which is prohibitive for practical applications. To this end, we introduce a novel Tensorized Graph Attention (TGA) mechanism, which leverages on the recent Graph Tensor Network (GTN) framework to efficiently process tensorized token embeddings via attention based graph filters. Such tensorized token embeddings are shown to effectively bypass the Curse of Dimensionality, reducing the parameter complexity of the attention mechanism from an exponential to a linear one in the embedding dimensions. The expressive power of the TGA framework is further enhanced by virtue of domain-aware graph convolution filters. Simulations across benchmark NLP paradigms verify the advantages of the proposed framework over existing attention models, at drastically lower parameter complexity.
Yao Lei Xu, Kriton Konstantinidis, Shengxi Li, Ljubisa Stankovic, Danilo P. Mandic
ICASSP5
2022 Hearables: Artefact removal in Ear-EEG for continuous 24/7 monitoring
abstract
Ear-worn devices offer the opportunity to measure vital signals in a 24/7 fashion, without the need of a clinician. These devices are however prone to motion artefacts, so that entire epochs of artefact-corrupt recordings are routinely discarded. This work aims at reducing the impact of artefacts introduced by a series of common real life daily activities such as talking, chewing, and walking while recording Electroencephalogram (EEG) from the ear canal. The approach used employs multiple external sensors, such as microphones and an accelerometer as means to capture the artefact. The proposed algorithm is a combination of Noise-Assisted Multivariate Empirical Mode Decomposition (NA-MEMD) with Adaptive Noise Cancellation (ANC), where each pair (EEG and motion sensors) of Intrinsic Mode Functions (IMFs) within NA-MEMD is fed independently to multiple Normalised Least Mean Square (NLMS) adaptive filters. The resulting denoised IMFs are then added up again to reconstruct the denoised EEG signal. Results across multiple subjects show that the so denoised EEG signals have reduced power in the frequency range occupied by artefacts. Also, different sensors provide different denoising performance in the tested artefacts, with the microphones being more sensitive to artefacts which cause internal motion within the ear-canal, such as chewing, and the accelerometer being more suitable for artefacts which come from full body movements of the subjects, such as walking.
Edoardo Occhipinti, Harry J. Davies, Ghena Hammour, Danilo P. Mandic
IJCNN4
2022 Fractional-Order Learning Systems
abstract
From the inaugural steps of McCulloch and Pitts to put forth a composition for an electrical brain, that combined with the conception of an adaptive leaning mechanism by Widrow and Hoff has given rise to the phenomena of intelligent machines, machine learning techniques have gained the status of a miracle solution in a myriad of scientific fields. At the heart of these techniques lies iterative optimisation processes that are derived based on first, and in some cases, second-order derivatives. This manuscript, however, aims to expand the mentioned framework to that of using fractional-order derivatives. The entire format of adaptation is revised form the perspective of fractional-order calculus and the appropriate framework for taking full advantage of the fractional-order calculus in learning and adaptation paradigms is formulated. For rigour, the structure of behavioural analysis and performance prediction of this novel class of learning machines is also forged.
Sayed Pouria Talebi, Stefan Werner 0001, Danilo P. Mandic
IJCNN3
2022 BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio Synthesis
abstract
Binaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warping of the mono audio, but also room reverberations and head/ear related filtration, which, however, are difficult to accurately simulate in traditional digital signal processing. In this paper, we formulate the synthesis process from a different perspective by decomposing the binaural audio into a common part that shared by the left and right channels as well as a specific part that differs in each channel. Accordingly, we propose BinauralGrad, a novel two-stage framework equipped with diffusion models to synthesize them respectively. Specifically, in the first stage, the common information of the binaural audio is generated with a single-channel diffusion model conditioned on the mono audio, based on which the binaural audio is generated by a two-channel diffusion model in the second stage. Combining this novel perspective of two-stage synthesis with advanced generative models (i.e., the diffusion models), the proposed BinauralGrad is able to generate accurate and high-fidelity binaural audio samples. Experiment results show that on a benchmark dataset, BinauralGrad outperforms the existing baselines by a large margin in terms of both object and subject evaluation metrics (Wave L2: $0.128$ vs. $0.157$, MOS: $3.80$ vs. $3.61$). The generated audio samples\footnote{\url{https://speechresearch.github.io/binauralgrad}} and code\footnote{\url{https://github.com/microsoft/NeuralSpeech/tree/master/BinauralGrad}} are available online.
Yichong Leng, Zehua Chen 0005, Junliang Guo, Haohe Liu, Jiawei Chen 0008, Xu Tan 0003, Danilo P. Mandic, Lei He 0005, Xiang-Yang Li 0001, Tao Qin 0001, Sheng Zhao 0002, Tie-Yan Liu
NeurIPS7
2022 Incremental deep learning for reflectivity data recognition in stomatology
abstract
Abstract The recognition of stomatological disorders and the classification of dental caries are important areas of biomedicine that can hugely benefit from machine learning tools for the construction of relevant mathematical models. This paper explores the possibility of using reflectivity data to distinguish between healthy tissues and caries by deep learning and multilayer convolutional neural networks. The experimental data set includes more than 700 observations recorded in the stomatology laboratory. For rigor, the results obtained from the deep learning systems are compared with those evaluated for selected sets of features estimated for each observation and classified by a decision tree, support vector machine (SVM), k-nearest neighbor, Bayesian methods, and two-layer neural networks. The classification accuracy obtained for the deep learning systems was 98.1% and 94.4% for data in the signal and spectral domains, respectively, in comparison with an accuracy of 97.2% and 87.2% evaluated by the SVM method. The proposed method conclusively demonstrates how the artificial intelligence and deep learning methodology can contribute to improved diagnosis of dental problem in stomatology.
Ales Procházka, Jindrich Charvát, Oldrich Vysata, Danilo P. Mandic
Neural Comput. Appl.4
2022 Online censoring based complex-valued adaptive filters
Engin Cemal Menguc, Min Xiang, Danilo P. Mandic
Signal Process.3
2022 A full second-order statistical analysis of strictly linear and widely linear estimators with MSE and Gaussian entropy criteria
Xing Zhang 0005, Yili Xia, Chunguo Li, Luxi Yang, Danilo P. Mandic
Signal Process.5
2022 Supervised Learning for Nonsequential Data: A Canonical Polyadic Decomposition Approach
Alexandros Haliassos, Kriton Konstantinidis, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2021 Robust PCA Through Maximum Correntropy Power Iterations
abstract
Principal component analysis (PCA) is considered a quintessential data analysis technique when it comes to describing linear relationships between the features of a dataset. However, the well-known lack of robustness of PCA for non-Gaussian data and/or outliers often makes its practical use unreliable. To this end, we introduce a robust formulation of PCA based on the maximum correntropy criterion (MCC). By virtue of MCC, robust operation is achieved by maximising the expected likelihood of Gaussian distributed reconstruction errors. The analysis shows that the proposed solution reduces to a generalised power iteration, whereby: (i) robust estimates of the principal components are obtained even in the presence of outliers; (ii) the number of principal components need not be specified in advance; and (iii) the entire set of principal components can be obtained, unlike existing approaches. The advantages of the proposed maximum correntropy power iteration (MCPI) are demonstrated through an intuitive numerical example.
Jean P. Chereau, Bruno Scalzo Dees, Danilo P. Mandic
ICASSP3
2021 Kernel Learning with Tensor Networks
abstract
The expressive power of Gaussian Processes (GPs) is largely attributed to their kernel function, which highlights the crucial role of kernel design. Efforts in this direction include modern Neural Network (NN) based kernel design, which despite success, suffers from the lack of interpretability and tendency to overfit. To this end, we introduce a Tensor Network (TN) approach to learning kernel embeddings, with a TN serving to map the input to a low dimensional manifold, where a suitable base kernel function can be applied. The proposed framework allows for joint learning of the TN and base kernel parameters using stochastic variational inference, while leveraging on the low-rank regularization and multi-linear nature of TNs to boost model performance and provide enhanced interpretability. Performance evaluation within the regression paradigm against TNs and Deep Kernels demonstrates the potential of the framework, providing conclusive evidence for promising future extensions to other learning paradigms.
Kriton Konstantinidis, Shengxi Li, Danilo P. Mandic
ICASSP3
2021 Nonstationary Portfolios: Diversification in the Spectral Domain
abstract
Classical portfolio optimization methods typically determine an optimal capital allocation through the implicit, yet critical, assumption of statistical time-invariance. Such models are inadequate for real-world markets as they employ standard time-averaging based estimators which suffer significant information loss if the market observables are non-stationary. To this end, we reformulate the portfolio optimization problem in the spectral domain to cater for the nonstationarity inherent to asset price movements and, in this way, allow for optimal capital allocations to be time-varying. Unlike existing spectral portfolio techniques, the proposed framework employs augmented complex statistics in order to exploit the interactions between the real and imaginary parts of the complex spectral variables, which in turn allows for the modelling of both harmonics and cyclostationarity in the time domain. The advantages of the proposed framework over traditional methods are demonstrated through numerical simulations using real-world price data.
Bruno Scalzo Dees, Alvaro Arroyo, Ljubisa Stankovic, Danilo P. Mandic
ICASSP4
2021 Graph Theory for Metro Traffic Modelling
abstract
A unifying graph theoretic framework for the modelling of metro transportation networks is proposed. This is achieved by first introducing a basic graph framework for the modelling of the London underground system from a diffusion law point of view. This forms a basis for the analysis of both station importance and their vulnerability, whereby the concept of graph vertex centrality plays a key role. We next explore k-edge augmentation of a graph topology, and illustrate its usefulness both for improving the network robustness and as a planning tool. Upon establishing the graph theoretic attributes of the underlying graph topology, we proceed to introduce models for processing data on such a metro graph. Commuter movement is shown to obey the Fick's law of diffusion, where the graph Laplacian provides an analytical model for the diffusion process of commuter population dynamics. Finally, we also explore the application of modern deep learning models, such as graph neural networks and hyper-graph neural networks, as general purpose models for the modelling and forecasting of underground data, especially in the context of the morning and evening rush hours. Comprehensive simulations including the passenger in- and outflows during the morning rush hour in London demonstrates the advantages of the graph models in metro planning and traffic management, a formal mathematical approach with wide economic implications.
Bruno Scalzo Dees, Yao Lei Xu, Anthony G. Constantinides, Danilo P. Mandic
IJCNN4
2021 Tensor-Train Recurrent Neural Networks for Interpretable Multi-Way Financial Forecasting
abstract
Recurrent Neural Networks (RNNs) represent the de facto standard machine learning tool for sequence modelling, owing to their expressive power and memory. However, when dealing with large dimensional data, the corresponding exponential increase in the number of parameters imposes a computational bottleneck. The necessity to equip RNNs with the ability to deal with the curse of dimensionality, such as through the parameter compression ability inherent to tensors, has led to the development of the Tensor-Train RNN (TT-RNN). Despite achieving promising results in many applications, the full potential of the TT-RNN is yet to be explored in the context of interpretable financial modelling, a notoriously challenging task characterized by multi-modal data with low signal-to-noise ratio. To address this issue, we investigate the potential of TT-RNN in the task of financial forecasting of currencies. We show, through the analysis of TT-factors, that the physical meaning underlying tensor decomposition, enables the TT-RNN model to aid the interpretability of results, thus mitigating the notorious “black-box” issue associated with neural networks. Furthermore, simulation results highlight the regularization power of TT decomposition, demonstrating the superior performance of TT-RNN over its uncompressed RNN counterpart and other tensor forecasting methods.
Yao Lei Xu, Giuseppe Giovanni Calvi, Danilo P. Mandic
IJCNN3
2021 Deep neural network representation and Generative Adversarial Learning
Ariel Ruiz-Garcia, Jürgen Schmidhuber, Vasile Palade, Clive Cheong Took, Danilo P. Mandic
Neural Networks5
2021 Convergence of the RMSProp deep learning method with penalty for nonconvex optimization
Dongpo Xu, Shengdong Zhang, Huisheng Zhang, Danilo P. Mandic
Neural Networks4
2021 A lower bound on the tensor rank based on its maximally square matrix unfolding
Giuseppe Giovanni Calvi, Bruno Scalzo Dees, Danilo P. Mandic
Signal Process.3
2021 A layer-wise distribution analysis of the WLMMSE-SIC MIMO receiver for rectilinear or quasi-rectilinear signals
Zhe Li 0007, Zhanyu Zhu, Danilo P. Mandic
Signal Process.4
2021 Online Censoring Based Weighted-Frequency Fourier Linear Combiner for Estimation of Pathological Hand Tremors
abstract
An online censoring (OC) based weighted-frequency Fourier linear combiner (OC-WFLC) adaptive filtering structure is proposed to reduce data processing costs in the estimation of pathological hand tremor (PHT) measurements. The proposed OC-WFLC is combined with the Fourier linear combiner (FLC) to effectively separate the PHT and voluntary movement from the hand tremor signal. The OC-WFLC is shown to adaptively extract the most informative frequency information, that is readily employed within the FLC to adaptively decompose the measurement signal into its PHT and voluntary movement components. The utilization of the OC strategy in the proposed framework is shown to significantly reduce data processing costs without adverse effects on the performance. Simulation results on real-world PHT data demonstrate the ability of the proposed OC-WFLC to yield a dramatic reduction of the processing time, a prerequisite for real-time rehabilitative, wearable, and assistive technology designed for PHT patients.
Engin Cemal Menguc, Salim Çinar, Min Xiang, Danilo P. Mandic
IEEE Signal Process. Lett.4
2021 Improved Coherence Index-Based Bound in Compressive Sensing
abstract
Within the compressive sensing (CS) paradigm, sparse signals can be reconstructed based on a reduced set of measurements, whereby reliability of the solution is determined by its uniqueness. With its mathematically tractable and feasible calculation, the coherence index is one of very few CS uniqueness metrics with considerable practical importance. We propose an improvement of the coherence-based uniqueness relation for the matching pursuit algorithms. Starting from a simple and intuitive derivation of the standard uniqueness condition, based on the coherence index, we derive a less conservative coherence index-based lower bound for signal sparsity. The results are generalized to the uniqueness condition of the l0-norm minimization for a signal represented in two orthonormal bases.
Ljubisa Stankovic, Milos Brajovic, Danilo P. Mandic, Isidora Stankovic, Milos Dakovic
IEEE Signal Process. Lett.3
2021 A Universal Framework for Learning the Elliptical Mixture Model
abstract
Mixture modeling using elliptical distributions promises enhanced robustness, flexibility, and stability over the widely employed Gaussian mixture model (GMM). However, existing studies based on the elliptical mixture model (EMM) are restricted to several specific types of elliptical probability density functions, which are not supported by general solutions or systematic analysis frameworks; this significantly limits the rigor in the design and power of EMMs in applications. To this end, we propose a novel general framework for estimating and analyzing the EMMs, achieved through the Riemannian manifold optimization. First, we investigate the relationships between Riemannian manifolds and elliptical distributions, and the so established connection between the original manifold and a reformulated one indicates a mismatch between these manifolds, a major cause of failure of the existing optimization for solving general EMMs. We next propose a universal solver that is based on the optimization of a redesigned cost and prove the existence of the same optimum as in the original problem; this is achieved in a simple, fast and stable way. We further calculate the influence functions of the EMM as theoretical bounds to quantify robustness to outliers. Comprehensive numerical results demonstrate the ability of the proposed framework to accommodate EMMs with different properties of individual functions in a stable way and with fast convergence speed. Finally, the enhanced robustness and flexibility of the proposed framework over the standard GMM are demonstrated both analytically and through comprehensive simulations.
Shengxi Li, Zeyang Yu, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2020 Solving General Elliptical Mixture Models through an Approximate Wasserstein Manifold
abstract
We address the estimation problem for general finite mixture models, with a particular focus on the elliptical mixture models (EMMs). Compared to the widely adopted Kullback–Leibler divergence, we show that the Wasserstein distance provides a more desirable optimisation space. We thus provide a stable solution to the EMMs that is both robust to initialisations and reaches a superior optimum by adaptively optimising along a manifold of an approximate Wasserstein distance. To this end, we first provide a unifying account of computable and identifiable EMMs, which serves as a basis to rigorously address the underpinning optimisation problem. Due to a probability constraint, solving this problem is extremely cumbersome and unstable, especially under the Wasserstein distance. To relieve this issue, we introduce an efficient optimisation method on a statistical manifold defined under an approximate Wasserstein distance, which allows for explicit metrics and computable operations, thus significantly stabilising and improving the EMM estimation. We further propose an adaptive method to accelerate the convergence. Experimental results demonstrate the excellent performance of the proposed EMM solver.
Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic
AAAI4
2020 Tensor Decompositions in Deep Learning
Davide Bacciu, Danilo P. Mandic
ESANN2
2020 Portfolio Cuts: A Graph-Theoretic Framework to Diversification
abstract
Investment returns naturally reside on irregular domains, however, standard multivariate portfolio optimization methods are agnostic to data structure. To this end, we investigate ways for domain knowledge to be conveniently incorporated into the analysis, by means of graphs. Next, to relax the assumption of the completeness of graph topology and to equip the graph model with practically relevant physical intuition, we introduce the portfolio cut paradigm. Such a graph-theoretic portfolio partitioning technique is shown to allow the investor to devise robust and tractable asset allocation schemes, by virtue of a rigorous graph framework for considering smaller, computationally feasible, and economically meaningful clusters of assets, based on graph cuts. In turn, this makes it possible to fully utilize the asset returns covariance matrix for constructing the portfolio, even without the requirement for its inversion. The advantages of the proposed framework over traditional methods are demonstrated through numerical simulations based on real-world price data.
Bruno Scalzo Dees, Ljubisa Stankovic, Anthony G. Constantinides, Danilo P. Mandic
ICASSP4
2020 A Low-Dimensionality Method for Data-Driven Graph Learning
abstract
In many graph signal processing applications, finding the topology of a graph is part of the overall data processing problem rather than a priori knowledge. Most of the approaches to graph topology learning are based on the assumption of graph Laplacian sparsity, with various additional constraints, followed by variations of the edge weights in the graph domain or the eigenvalues in the graph spectral domain. These domains are high-dimensional, since their dimension is at least equal to the order of the number of vertices. In this paper, we propose a numerically efficient method for estimating of the normalized Laplacian through its eigenvalues estimation and by promoting its sparsity. The minimization problem is solved in quite a low-dimensional space, related to the polynomial order of the underlying system on a graph corresponding to the the observed data. The accuracy of the results is tested on numerical example.
Ljubisa Stankovic, Milos Dakovic, Danilo P. Mandic, Milos Brajovic, Bruno Scalzo Dees, Anthony G. Constantinides
ICASSP3
2020 A Probabilistic Beat-to-Beat Filtering Model for Continuous and Accurate Blood Pressure Estimation
abstract
Cuffless technologies provide a convenient platform for remote and continuous blood pressure (BP) monitoring, however, signal recordings employed for cuffless BP estimation, which are based on the electrocardiogram (ECG) and photo-plethysmography (PPG) signals, are frequently corrupted with measurement noise and artefacts. Consequently, even a small portion of abnormal data samples can severely impact the overall signal quality and therefore lead to significantly distorted BP value estimates. To this end, a data-driven model is proposed to infer the beat-to-beat signal quality for the ECG, PPG and BP signal recordings, whereby high-quality and low-quality (outlier) beats are detected using a probabilistic model chosen according to the maximum entropy principle. Physiological rules are also imposed to guarantee that each filtered sample is physiologically meaningful. The advantages of the proposed filtering framework for both systolic blood pressure and diastolic blood pressure estimation are demonstrated through the analysis and estimation of 12,000 clinical BP recordings, consisting of over 200,000 test samples.
Zehua Chen 0005, Bruno Scalzo Dees, Danilo P. Mandic
IJCNN3
2020 Reciprocal Adversarial Learning via Characteristic Functions
abstract
Generative adversarial nets (GANs) have become a preferred tool for tasks involving complicated distributions. To stabilise the training and reduce the mode collapse of GANs, one of their main variants employs the integral probability metric (IPM) as the loss function. This provides extensive IPM-GANs with theoretical support for basically comparing moments in an embedded domain of the \textit{critic}. We generalise this by comparing the distributions rather than their moments via a powerful tool, i.e., the characteristic function (CF), which uniquely and universally comprising all the information about a distribution. For rigour, we first establish the physical meaning of the phase and amplitude in CF, and show that this provides a feasible way of balancing the accuracy and diversity of generation. We then develop an efficient sampling strategy to calculate the CFs. Within this framework, we further prove an equivalence between the embedded and data domains when a reciprocal exists, where we naturally develop the GAN in an auto-encoder structure, in a way of comparing everything in the embedded space (a semantically meaningful manifold). This efficient structure uses only two modules, together with a simple training strategy, to achieve bi-directionally generating clear images, which is referred to as the reciprocal CF GAN (RCF-GAN). Experimental results demonstrate the superior performances of the proposed RCF-GAN in terms of both generation and reconstruction.
Shengxi Li, Zeyang Yu, Min Xiang, Danilo P. Mandic
NeurIPS4
2020 SINR Analysis Of Mimo Systems With Widely Linear MMSE Receivers For The Reception Of Real-Valued Constellations
abstract
Although the widely linear minimum mean-square error (WLMMSE) receiver has been widely applied in multiple-input-multiple-output (MIMO) systems, there has been no theoretical analysis to quantify its distribution of signal-to-interference-plus-noise ratio (SINR) in arbitrary fading environments. In this paper, the closed-from expression of SINR at the output of the WLMMSE detection is presented, which, in its essence, can be interpreted as the sum of a series of gamma distributed random variables. The general probability density function of SINR is derived for the first time, which is explicitly expressed in terms of the confluent Lauricella hypergeometric function. Simulations on MIMO transmission systems over Rayleigh fading channels support the analytic results.
Zhe Li 0007, Wenjiang Pei, Yili Xia, Danilo P. Mandic
PIMRC6
2020 Improperness Based SINR Analysis of GFDM Systems Under Joint Tx and Rx I/Q Imbalance
abstract
Adverse impacts of in-phase and quadrature-phase (I/Q) imbalance in both the transmitter (Tx) and receiver (Rx) are quantified for the generalized frequency division multiplexing (GFDM) based transmission over frequency selective fading channels. To this end, we first equip the standard signal-to-interference-plus-noise (SINR) performance evaluation with the ability to consider second-order noncircular (improper) signals, and thus precisely evaluate performance deterioration caused by I/Q distortions over the in-phase (I) and quadrature-phase (Q) channels of a transmission system. Next, we propose a novel means to evaluate the individual SINR contributions from both the channels of GFDM, and hence, provide more meaningful insights into the underlying wireless transmission in the presence of complex non-circularity. This is accompanied by an account of complete augmented second-order statistics of I/Q imbalanced GFDM waveforms which caters for various sources of complex improperness. Simulations in the GFDM system setting support our analysis.
Hao Cheng 0006, Yili Xia, Yongming Huang 0001, Luxi Yang, Zixiang Xiong, Danilo P. Mandic
WCNC6
2020 Quadratic programming over ellipsoids with applications to constrained linear regression and tensor decomposition
Anh Huy Phan 0001, Masao Yamagishi, Danilo P. Mandic, Andrzej Cichocki
Neural Comput. Appl.3
2020 On the decomposition of multichannel nonstationary multicomponent signals
Ljubisa Stankovic, Milos Brajovic, Milos Dakovic, Danilo P. Mandic
Signal Process.4
2020 Fractional-Order Correntropy Adaptive Filters for Distributed Processing of $\alpha$-Stable Signals
abstract
This work revisits the problem of distributed adaptive filtering in multi-agent sensor networks. In contrast to classical approaches, the formulation relaxes the Gaussian assumption on the signal and noise to the generalized setting of α-stable distributions that do not possess second- and higher-order statistical moments. Most importantly, the considered scenario allows for different characteristic exponents throughout the network. Drawing upon ideas from correntropy-type local similarity measures and fractional-order calculus, a novel class of distributed fractional-order correntropy adaptive filters, that are robust against the jittery behavior of α-stable signals, is derived and their convergence criterion is established. The effectiveness of the proposed algorithms, as compared to existing distributed adaptive filtering techniques, is demonstrated via simulation examples.
Vinay Chakravarthi Gogineni, Sayed Pouria Talebi, Stefan Werner 0001, Danilo P. Mandic
IEEE Signal Process. Lett.4
2020 Tensor Networks for Latent Variable Analysis: Higher Order Canonical Polyadic Decomposition
abstract
The canonical polyadic decomposition (CPD) is a convenient and intuitive tool for tensor factorization; however, for higher order tensors, it often exhibits high computational cost and permutation of tensor entries, and these undesirable effects grow exponentially with the tensor order. Prior compression of tensor in-hand can reduce the computational cost of CPD, but this is only applicable when the rank R of the decomposition does not exceed the tensor dimensions. To resolve these issues, we present a novel method for CPD of higher order tensors, which rests upon a simple tensor network of representative inter-connected core tensors of orders not higher than 3. For rigor, we develop an exact conversion scheme from the core tensors to the factor matrices in CPD and an iterative algorithm of low complexity to estimate these factor matrices for the inexact case. Comprehensive simulations over a variety of scenarios support the proposed approach.
Anh Huy Phan 0001, Andrzej Cichocki, Ivan V. Oseledets, Giuseppe Giovanni Calvi, Salman Ahmadi-Asl, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.6
2020 Tensor Networks for Latent Variable Analysis: Novel Algorithms for Tensor Train Approximation
abstract
Decompositions of tensors into factor matrices, which interact through a core tensor, have found numerous applications in signal processing and machine learning. A more general tensor model that represents data as an ordered network of subtensors of order-2 or order-3 has, so far, not been widely considered in these fields, although this so-called tensor network (TN) decomposition has been long studied in quantum physics and scientific computing. In this article, we present novel algorithms and applications of TN decompositions, with a particular focus on the tensor train (TT) decomposition and its variants. The novel algorithms developed for the TT decomposition update, in an alternating way, one or several core tensors at each iteration and exhibit enhanced mathematical tractability and scalability for large-scale data tensors. For rigor, the cases of the given ranks, given approximation error, and the given error bound are all considered. The proposed algorithms provide well-balanced TT-decompositions and are tested in the classic paradigms of blind source separation from a single mixture, denoising, and feature extraction, achieving superior performance over the widely used truncated algorithms for TT decomposition.
Anh Huy Phan 0001, Andrzej Cichocki, André Uschmajew, Petr Tichavský, George Luta, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.6
2020 Complex Properness Inspired Blind Adaptive Frequency-Dependent I/Q Imbalance Compensation for Wideband Direct-Conversion Receivers
abstract
Direct-conversion receivers (DCRs) have been adopted in wideband communication systems owing to their simple structure and low cost, however, their operation is affected by amplitude and phase mismatches between their analog inphase (I) and quadrature (Q) branches, as well as the discrepancy of low-pass filter coefficients between these two channels. In this paper, a blind adaptive frequency-dependent I/Q imbalance compensator is proposed, which exploits the complex properness (second-order circularity) of ideal constellation mappings to provide more enhanced insight into the problem setting within the proposed compensator. This serves as a basis for a novel full second-order performance assessment framework, which is established through a joint consideration of the weight error covariance and complementary covariance in both the transient and steady-state stages. This conjoint analysis is further shown to facilitate accurate quantification of the overall mirror-frequency interference attenuation capability of the proposed compensator. Simulation results in an orthogonal frequency division multiplexing (OFDM) transmission system demonstrate the excellent performance of the proposed compensator.
Xing Zhang 0005, Yili Xia, Chunguo Li, Luxi Yang, Danilo P. Mandic
IEEE Trans. Wirel. Commun.5
2019 Tensor Ring Decomposition with Rank Minimization on Latent Space: An Efficient Approach for Tensor Completion
abstract
In tensor completion tasks, the traditional low-rank tensor decomposition models suffer from the laborious model selection problem due to their high model sensitivity. In particular, for tensor ring (TR) decomposition, the number of model possibilities grows exponentially with the tensor order, which makes it rather challenging to find the optimal TR decomposition. In this paper, by exploiting the low-rank structure of the TR latent space, we propose a novel tensor completion method which is robust to model selection. In contrast to imposing the low-rank constraint on the data space, we introduce nuclear norm regularization on the latent TR factors, resulting in the optimization step using singular value decomposition (SVD) being performed at a much smaller scale. By leveraging the alternating direction method of multipliers (ADMM) scheme, the latent TR factors with optimal rank and the recovered tensor can be obtained simultaneously. Our proposed algorithm is shown to effectively alleviate the burden of TR-rank selection, thereby greatly reducing the computational cost. The extensive experimental results on both synthetic and real-world data demonstrate the superior performance and efficiency of the proposed approach against the state-of-the-art algorithms.
Longhao Yuan, Chao Li 0013, Danilo P. Mandic, Jianting Cao, Qibin Zhao
AAAI3
2019 Widely Linear Complex-Valued Autoencoder: Dealing with Noncircularity in Generative-Discriminative Models
Zeyang Yu, Shengxi Li, Danilo P. Mandic
ICANN (1)3
2019 Support Tensor Machine for Financial Forecasting
abstract
Past decades have witnessed excessive use of the Support Vector Machines (SVMs) in financial contexts. Despite their success, given the inherently multivariate nature of financial indices, the vector-based nature of SVM will inevitably lead to a loss of information, owing to its inability to fully exploit the available multi-way data structure. Motivated by the superior structural information content in tensors over vectors, we investigate the usefulness of the tensor extension of SVM, termed the Support Tensor Machine (STM), in the financial application of forecasting the daily direction of movement of price of the S&P 500 financial index. A computationally efficient least-squares formulation of STM (LS-STM) is considered, and the tensorized data includes the VIX, i.e. the implied index of volatility of the S&P 500, the GC1 gold commodity, and the S&P 500 itself. Next, a method to allow kernel usage in LS-STM is also introduced. The LS-STM based results are interpreted probabilistically via Platt scaling, while performance is evaluated comprehensively in terms of accuracy rate, annualized Sharpe ratio, and annualized volatility. All performance metrics conclusively demonstrate the superiority of LS-STM over standard SVM.
Giuseppe Giovanni Calvi, Vladimir Lucic, Danilo P. Mandic
ICASSP3
2019 Smart DSP for a Smarter Power Grid: Teaching Power System Analysis through Signal Processing
abstract
The future Smart Grid represents an extraordinary opportunity to transform the ways we currently approach energy into a new era of low-carbon, renewable, and efficient solutions which will ultimately have a significant impact on both the environment and economy. This effort requires close collaboration of experts from the Power, Digital Signal Processing (DSP) and Machine Learning (ML) communities, with the common language between these diverse disciplines an important first step in this endeavour. To promote seamless transition of ideas, we here establish a duality between the Clarke transform, a workhorse in Power Grid analysis, and principal component analysis (PCA), a staple subspace method in DSP/ML. Upon highlighting the limitations of the Clarke transform in off-nominal unbalanced power system conditions, we illuminate a DSP-enabled class of self-balancing solutions, referred to as the Smart Transforms, based on adaptive complex widely linear modelling. Examples on system frequency estimation support the approach.
Ahmad Moniri, Anthony G. Constantinides, Danilo P. Mandic
ICASSP3
2019 Tracking Dynamic Systems in α-Stable Environments
abstract
In order to accommodate for modern adaptive filtering applications, the classic adaptive filtering paradigm is considered from a more general perspective. The new formulation allows for time dependent variations in the state of the system and more importantly it relaxes the Gaussian assumption to the generalized setting of α-stable distributions. In this work, based on the principles of gradient descent and fractional-order calculus, a cost-effective technique for tracking the state of such a system is derived. For rigour, performance of the derived filtering technique is analyzed and convergence conditions are established.
Sayed Pouria Talebi, Stefan Werner 0001, Shengxi Li, Danilo P. Mandic
ICASSP4
2019 Quaternion-Valued Adaptive Filtering via Nesterov's Extrapolation
abstract
A new quaternion-valued adaptive filtering algorithm based on extrapolated weight methods is proposed. The proposed algorithm belongs to the class of conjugate direction algorithms [1]. This class of extrapolation (momentum) based algorithms is preferred to RLS-based algorithms when the matrix inversion should be avoided, e.g. in the case of non-vector signals, sparse signals or non-stationary signals. This paper introduces Nesterov's optimal gradient methods in widely linear quaternion adaptive filtering. The resulting class of algorithm is shown to both have similar computational complexity and comparable performance to WLQRLS; however, the proposed method is more stable and outperforms WLQRLS in the non-stationary case.
Thiernithi Variddhisaï, Min Xiang, Scott C. Douglas, Danilo P. Mandic
ICASSP4
2019 Simultaneous DFT and IDFT through Widely Linear CLMS
abstract
Complex least mean square (CLMS) based adaptive computation of discrete orthogonal transforms has been extensively investigated in the literature. However, all of these results provide only a means for the calculation of either forward orthogonal transforms or their inverse orthogonal transforms, separately. In this work, a way to simultaneously calculate the discrete Fourier transform (DFT) and the inverse DFT (IDFT) is established via the widely linear (WL) signal processing framework. We show that by appropriately selecting the input vector and adaptation speed of the widely linear complex least mean square (WL-CLMS), the resulting spectrum analyzer is capable of simultaneously performing DFT and IDFT of the signal to be Fourier analyzed in both the block-based and online manners.
Xing Zhang 0005, Bruno Scalzo Dees, Chunguo Li, Yili Xia, Luxi Yang, Danilo P. Mandic
ICASSP6
2019 A cost-effective nonlinear self-interference canceller in full-duplex direct-conversion transceivers
Zhe Li 0007, Yili Xia, Wenjiang Pei, Danilo P. Mandic
Signal Process.4
2019 A class of multidimensional NIPALS algorithms for quaternion and tensor partial least squares regression
Alexander Stott, Bruno Scalzo Dees, Ilia Kisil, Danilo P. Mandic
Signal Process.4
2019 Complex-Valued Nonlinear Adaptive Filters With Applications in $\alpha$-Stable Environments
abstract
A nonlinear adaptive filtering framework for processing complex-valued signals is derived. The introduced adaptive filter extends the fractional-order framework of the authors for dealing with real-valued signals to the complex domain via the augmented statistical approach to complex-valued signal processing. This results in a versatile class of adaptive filtering techniques, which allows the classical Gaussian assumption to be extended to the generalized context of α-stables. For rigor, the performance of the introduced adaptive filtering framework is analyzed, its convergence criteria is established, and its application in tracking signals of chaotic systems is demonstrated using simulations.
Sayed Pouria Talebi, Stefan Werner 0001, Danilo P. Mandic
IEEE Signal Process. Lett.3
2019 Complementary Cost Functions for Complex and Quaternion Widely Linear Estimation
abstract
Widely linear (WL) models have been demonstrated to be superior to conventional strictly linear models for the estimation of noncircular complex and quaternion signals. Existing studies on their performance bounds focus on the analysis of mean square error (MSE). However, the single degree of freedom within standard MSE allows for only the minimization of error power, with no means to understand how the error contribution is distributed across the data channels. To this end, we introduce novel complex and quaternion valued complementary quadratic cost functions for complex and quaternion signal estimation, which are extensions of the recently proposed complementary MSE metric. It is shown that for WL minimum MSE estimation and least squares regression, the complementary cost function and the standard cost function attain the same stationary point. We also show that for the former, the stationary point is a saddle point. This novel finding provides insight into the performance of complex and quaternion WL estimators, and offers a rigorous foundation for further developments in this field.
Min Xiang, Yili Xia, Danilo P. Mandic
IEEE Signal Process. Lett.3
2019 Multiple-Model Adaptive Estimation for 3-D and 4-D Signals: A Widely Linear Quaternion Approach
abstract
Quaternion state estimation techniques have been used in various applications, yet they are only suitable for dynamical systems represented by a single known model. In order to deal with model uncertainty, this paper proposes a class of widely linear quaternion multiple-model adaptive estimation (WL-QMMAE) algorithms based on widely linear quaternion Kalman filters and Bayesian inference. The augmented second-order quaternion statistics is employed to capture complete second-order statistical information in improper quaternion signals. Within the WL-QMMAE framework, a widely linear quaternion interacting multiple-model algorithm is proposed to track time-variant model uncertainty, while a widely linear quaternion static multiple-model algorithm is proposed for time-invariant model uncertainty. A performance analysis of the proposed algorithms shows that, as expected, the WL-QMMAE reduces to semiwidely linear QMMAE for [Formula: see text]-improper signals and further reduces to strictly linear QMMAE for proper signals. Simulation results indicate that for improper signals, the proposed WL-QMMAE algorithms exhibit an enhanced performance over their strictly linear counterparts. The effectiveness of the proposed recursive performance analysis algorithm is also validated.
Min Xiang, Bruno Scalzo Dees, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2019 Joint Channel Estimation and Tx/Rx I/Q Imbalance Compensation for GFDM Systems
abstract
Generalized frequency division multiplexing (GFDM) has become one of the most important waveform candidates for beyond 5G (B5G) communications. However, physical distortions, such as in-phase and quadrature (I/Q) imbalance caused by the imperfections of radio frequency (RF) components within direct-conversion transceivers (DCTs), may cause severe performance degradation in GFDM-based wireless systems. To this end, we first conduct a rigorous sum rate analysis to quantify the impact of I/Q imbalance in both the transmitter and the receiver on the GFDM wireless transmission. An efficient I/Q imbalance compensation scheme is next proposed based on pilots; this is achieved through a nonlinear least squares analysis of the joint channel and I/Q imbalance estimation, and a simple symbol detection procedure. For rigor, the Cramer-Rao lower bounds for both the I/Q imbalance parameters and the channel coefficients are also derived. The simulation results illustrate that the mean square error performance of the proposed estimator closely approaches the corresponding CRLB over static frequency selective channels, thus significantly reducing the sensitivity of GFDM DCTs to physical I/Q impairments.
Hao Cheng 0006, Yili Xia, Yongming Huang 0001, Luxi Yang, Danilo P. Mandic
IEEE Trans. Wirel. Commun.5
2018 Hypergraph p-Laplacian: A Differential Geometry View
abstract
The graph Laplacian plays key roles in information processing of relational data, and has analogies with the Laplacian in differential geometry. In this paper, we generalize the analogy between graph Laplacian and differential geometry to the hypergraph setting, and propose a novel hypergraph p-Laplacian. Unlike the existing two-node graph Laplacians, this generalization makes it possible to analyze hypergraphs, where the edges are allowed to connect any number of nodes. Moreover, we propose a semi-supervised learning method based on the proposed hypergraph p-Laplacian, and formalize them as the analogue to the Dirichlet problem, which often appears in physics. We further explore theoretical connections to normalized hypergraph cut on a hypergraph, and propose normalized cut corresponding to hypergraph p-Laplacian. The proposed p-Laplacian is shown to outperform standard hypergraph Laplacians in the experiment on a hypergraph semi-supervised learning and normalized cut setting.
Shota Saito, Danilo P. Mandic, Hideyuki Suzuki
AAAI2
2018 Widely Linear CLMS Based Cancelation of Nonlinear Self -Interference in Full-Duplex Direct-Conversion Transceivers
abstract
An augmented nonlinear complex LMS (ANCLMS) algorithm is proposed to adaptively mitigate both the linear and nonlinear self-interference (SI) components in a full-duplex direct-conversion transceiver (DCT). A data prewhitening scheme, which exploits the known SI signal distributions, is also adopted to accelerate the convergence. Theoretical mean and mean square performance evaluations of the proposed SI canceller are performed and fully support the proposed approach. Computer simulations on wireless local area network (WLAN) standard compliant waveforms in practical full-duplex (FD) direct-conversion transceiver settings support the analysis.
Zhe Li 0007, Wenjiang Pei, Yili Xia, Kai Wang 0020, Danilo P. Mandic
ICASSP5
2018 Complementary Complex-Valued Spectrum for Real-Valued Data: Real Time Estimation of the Panorama Through Circularity-Preserving Dft
abstract
This work sheds a new light on the spectral whitening effects of the sliding discrete Fourier transform (DFT) and uses it as a basis for a novel technique for circularity-preserving spectral estimation. This makes it possible to utilise full available spectral information, unlike the existing methods which ignore the phase spectrum. We then use the so introduced circularity-preserving DFT to show that the Wiener filter can be used to estimate the recently introduced second-order complementary spectral measure, termed the panorama, even in the critical cases of short data windows and incoherent sampling. Numerical examples demonstrate the ability of the proposed procedure to estimate the spectral circularity and the panorama, even in a streaming-data setting for which the current methods are inadequate.
Bruno Scalzo Dees, Scott C. Douglas, Danilo P. Mandic
ICASSP3
2018 Correntropy-Based Adaptive Filtering of Noncircular Complex Data
abstract
Real world complex-valued signals typically exhibit rotation-dependent distributions (noncircularity), and significant performance gains in learning algorithms can be obtained by accounting for information beyond the standard second-order noncircularity (impropriety). To this end, we introduce a new closed form definition of complex correntropy which is general enough to cater for both circular and noncircular distributions in complex data, and serves as a basis for a novel cost function for widely linear adaptive filtering, termed the maximum improper complex corren-tropy criterion (MICCC). A stochastic gradient adaptive filtering algorithm is developed based on the MICCC, and its standard and complementary convergence and stability analyses are conducted with respect to both the circularity of the estimation error and the kernel size in the underlying Parzen estimator. Performance advantages over the strictly linear correntropy algorithm (MCCC) and the mean square error based complex least mean square (CLMS) and augmented CLMS (ACLMS) are demonstrated through analysis and simulations.
Bruno Scalzo Dees, Yili Xia, Scott C. Douglas, Danilo P. Mandic
ICASSP4
2018 Affine-Projection Least-Mean-Magnitude-Phase Algorithms Using a Posteriori Updates
abstract
The least-mean-magnitude-phase (LMMP) algorithm is useful for complex-valued signal processing applications where precise control of magnitude and/or phase error information can provide improved estimation performance. Because it is a gradient procedure, however, the convergence speed of the algorithm can be limited for correlated input signals. In this paper, we derive affine-projection least-mean-magnitude-phase (AP-LMMP) algorithms based on an a posteriori update relation that have improved convergence performance over that of the LMMP algorithm without significant increases in complexity. We employ different nonlinear lookahead approaches depending on the projection order to compute the magnitudes of the a posteriori output signals and use these to implement the coefficient updates. Simulations indicate that AP-LMMP algorithms can outperform other algorithms in situations where their use is appropriate.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2018 EAR-EEG for Detecting Inter-Brain Synchronisation in Continuous Cooperative Multi-Person Scenarios
abstract
The hyperscanning method simultaneously acquires and relates cerebral data from two participants while performing cooperative activities. The aim of this work is to evaluate the performance of our novel EEG recording concept, termed ear-EEG, against on-scalp EEG as an alternative, user-friendly data acquisition approach for hyperscanning, in the task of identifying the most robust, EEG subbands for inter-individual neuronal synchrony detection in cooperative multi-player gaming. This is achieved through the estimation of neuronal synchrony produced by a highly localised time-frequency data association measure, termed intrinsic synchrosqueezing coherence (ISC). It is shown that for both the recording modalities the lower theta band is the most robust neuronal marker for inter-brain synchronisation during our own cooperative game called Bar Balancing. This is because the lower theta band yields: (i) highest correlation in neuronal synchrony between the modalities; and (ii) enhanced discrimination ability of significant neuronal synchrony detected in the easy and hard tasks in both the modalities.
Apit Hemakom, Valentin Goverdovsky, Danilo P. Mandic
ICASSP3
2018 Common and Individual Feature Extraction Using Tensor Decompositions: a Remedy for the Curse of Dimensionality?
abstract
A novel method for common and individual feature analysis from exceedingly large-scale data is proposed, in order to ensure the tractability of both the computation and storage and thus mitigate the curse of dimensionality, a major bottleneck in modern data science. This is achieved by making use of the inherent redundancy in so-called multi-block data structures, which represent multiple observations of the same phenomenon taken at different times, angles or recording conditions. Upon providing an intrinsic link between the properties of the outer vector product and extracted features in tensor decompositions (TDs), the proposed common and individual information extraction from multi-block data is performed through constraints which impose physical meaning on otherwise unconstrained factorisation approaches. This is shown to dramatically reduce the dimensionality of search spaces in subsequent classification procedures and to yield greatly enhanced accuracy. Simulations on a multi-class classification task of large-scale extraction of individual features from a collection of partially related real-world images demonstrate the advantages of the “blessing of dimensionality” associated with TDs.
Ilia Kisil, Giuseppe Giovanni Calvi, Andrzej Cichocki, Danilo P. Mandic
ICASSP4
2018 Automatic detection of drowsiness using in-ear EEG
abstract
Sleep monitoring with wearable electroencephalography (EEG) has recently been validated and reported in the research community. One such device is our ultra-wearable, unobtrusive, and inconspicuous in-ear EEG system, which has already been demonstrated to be next-generation solution for out-of-clinic sleep monitoring. We here provide a further proof of concept of the utility of ear-EEG in day time drowsiness monitoring in the real-world. For rigour, hypnograms are obtained from manually scored daytime nap recordings from twentythree subjects, while a complexity science feature-structural complexity extracted from scalp- and ear-EEG recordings - is used in the classification stage, in conjunction with a binary-class support vector machine (SVM). The achieved drowsiness classification accuracies range from 80.0% to 82.9% for ear-EEG, with the corresponding accuracies for scalp-EEG ranging from 86.8 % to 88.8 %. Given the notoriously difficult to classify drowsiness related changes in EEG (similar to the issues with the NREM Stage 1), this conclusively confirms the feasibility of in-ear EEG for automatic light sleep classification. This also promises a key stepping stone towards continuous, discreet, and user-friendly wearable out-of-clinic drowsiness monitoring in the real-world, with numerous applications in the monitoring the state of body and mind of pilots, train drivers, and tele-operators.
Yousef Alqurashi, Mary J. Morrell, Danilo P. Mandic
IJCNN4
2018 Probabilistic guidance for catheter tip motion in cardiac ablation procedures
Mihaela Constantinescu, Su-Lin Lee, Sabine Ernst, Apit Hemakom, Danilo P. Mandic, Guang-Zhong Yang
Medical Image Anal.5
2018 Time-frequency decomposition of multivariate multicomponent signals
Ljubisa Stankovic, Danilo P. Mandic, Milos Dakovic, Milos Brajovic
Signal Process.2
2018 Widely linear complex partial least squares for latent subspace regression
Alexander Stott, Sithan Kanna, Danilo P. Mandic
Signal Process.3
2018 Performance analysis of the deficient length augmented CLMS algorithm for second order noncircular complex signals
Yili Xia, Scott C. Douglas, Danilo P. Mandic
Signal Process.3
2018 A perspective on CLMS as a deficient length augmented CLMS: Dealing with second order noncircularity
Yili Xia, Scott C. Douglas, Danilo P. Mandic
Signal Process.3
2018 Simultaneous diagonalisation of the covariance and complementary covariance matrices in quaternion widely linear signal processing
Min Xiang, Shirin Enshaeifar, Alexander Stott, Clive Cheong Took, Yili Xia, Sithan Kanna, Danilo P. Mandic
Signal Process.7
2018 Distributed Adaptive Filtering of α-Stable Signals
abstract
A cost-effective framework for distributed adaptive filtering of α-stable signals over sensor networks is proposed. First, the filtering paradigm of α-stable signals through multiple observations made over a network of sensors is revisited and an optimal solution is formulated. Then, an adaptive gradient descent based algorithm for distributed real-time filtering of α-stable signals via multiagent networks is derived. This not only provides an approximation of the formulated optimal solution, but also a cost-effective algorithm that scales with the size of the network. Moreover, performance of the derived algorithm is analyzed and convergence conditions are established.
Sayed Pouria Talebi, Stefan Werner 0001, Danilo P. Mandic
IEEE Signal Process. Lett.3
2018 In-Ear EEG Biometrics for Feasible and Readily Collectable Real-World Person Authentication
abstract
The use of electroencephalogram (EEG) as a biometrics modality has been investigated for about a decade; however, its feasibility in real-world applications is not yet conclusively established, mainly due to the issues with collectability and reproducibility. To this end, we propose a readily deployable EEG biometrics system based on a “one-fits-all” viscoelastic generic in-ear EEG sensor (collectability), which does not require skilled assistance or cumbersome preparation. Unlike most existing studies, we consider data recorded over multiple recording days and for multiple subjects (reproducibility) while, for rigour, the training and test segments are not taken from the same recording days. A robust approach is considered based on the resting state with eyes closed paradigm, the use of both parametric (autoregressive model) and non-parametric (spectral) features, and supported by simple and fast cosine distance, linear discriminant analysis, and support vector machine classifiers. Both the verification and identification forensics scenarios are considered and the achieved results are on par with the studies based on impractical on-scalp recordings. Comprehensive analysis over a number of subjects, setups, and analysis features demonstrates the feasibility of the proposed ear-EEG biometrics, and its potential in resolving the critical collectability, robustness, and reproducibility issues associated with current EEG biometrics.
Valentin Goverdovsky, Danilo P. Mandic
IEEE Trans. Inf. Forensics Secur.3
2017 Single-channel Wiener filtering of deterministic signals in stochastic noise using the panorama
abstract
The Wiener filter is a well-known signal processing method for improving a noisy signal's quality. The Wiener filter requires either knowledge of or estimates of the power spectra of the signal-of-interest and of the undesired noise, leading to implementation challenges. In this paper, we show how a recently-developed second-order signal quantity termed the panorama can be employed to compute the Wiener filter for deterministic signals - containing nearly-constant frequency and phase components - in additive stochastic noise. We first show how the Wiener filter transfer function is related to the absolute value of the panorama and the power spectrum of the noisy measured signal. We then provide a practical procedure for estimating the absolute value of the panorama across frequency for single-channel sampled recordings. Numerical examples show the ability of the proposed procedure for estimating the panorama and for computing the Wiener filter for noisy deterministic signals.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2017 An online NIPALS algorithm for Partial Least Squares
abstract
Partial Least Squares (PLS) has been gaining popularity as a multivariate data analysis tool due to its ability to cater for noisy, collinear and incomplete data-sets. However, most PLS solutions are designed as block-based algorithms, rendering them unsuitable for environments with streaming data and non-stationary statistics. To this end, we propose an online version of the nonlinear iterative PLS (NIPALS) algorithm, based on a recursive computation of covariance matrices and gradient-based techniques to compute eigenvectors of the relevant matrices. Simulations over synthetic data show that the regression coefficients from the proposed online PLS algorithm converge to those of its block-based counterparts.
Alexander Stott, Sithan Kanna, Danilo P. Mandic, William T. Pike
ICASSP3
2017 Cost-effective diffusion Kalman filtering with implicit measurement exchanges
abstract
A resource effective extension to the class of distributed real-time diffusion Kalman filters is proposed. The proposed scheme removes the need to share measurement variables explicitly, by sharing only the state estimates and state error covariance matrices which implicitly contain the information about the measurements, observations matrices, and noise covariance matrices. The proposed distributed Kalman filter is quiet general, and its performance is comparable to that of existing diffusion based schemes, while having lower communication and computational requirements per-iteration compared to current distributed Kalman filtering algorithms.
Sayed Pouria Talebi, Sithan Kanna, Yili Xia, Danilo P. Mandic
ICASSP4
2017 Complexity science for sleep stage classification from EEG
abstract
Automatic sleep stage classification is an important paradigm in computational intelligence and promises considerable advantages to the health care. Most current automated methods require the multiple electroencephalogram (EEG) channels and typically cannot distinguish the S1 sleep stage from EEG. The aim of this study is to revisit automatic sleep stage classification from EEGs using complexity science methods. The proposed method applies fuzzy entropy and permutation entropy as kernels of multi-scale entropy analysis. To account for sleep transition, the preceding and following 30 seconds of epoch data were used for analysis as well as the current epoch. Combining the entropy and spectral edge frequency features extracted from one EEG channel, a multi-class support vector machine (SVM) was able to classify 93.8% of 5 sleep stages for the SleepEDF database [expanded], with the sensitivity of S1 stage was 49.1%. Also, the Kappa's coefficient yielded 0.90, which indicates almost perfect agreement.
Tricia Adjei, Yousef Alqurashi, David Looney, Mary J. Morrell, Danilo P. Mandic
IJCNN6
2017 Translation invariant multi-scale signal denoising based on goodness-of-fit tests
Naveed ur Rehman, Syed Zain Abbas, Anum Asif, Anum Javed, Khuram Naveed, Danilo P. Mandic
Signal Process.6
2017 Cost-effective quaternion minimum mean square error estimation: From widely linear to four-channel processing
Min Xiang, Clive Cheong Took, Danilo P. Mandic
Signal Process.3
2017 Distributed Particle Filtering of α-Stable Signals
abstract
In order to present an inclusive framework for distributed estimation/tracking of α-stable signals, a novel distributed particle filtering algorithm is developed. This is achieved through the reformulation of the particle filtering operations from the point of view of the characteristic function and is based on the decomposition of the operations of the particle filter so that they can be distributed among the agents of a sensor network while allowing each agent to retain an estimate of the state vector. In contrast to current distributed particle filtering techniques that approximate distributions with Gaussian mixtures through empirical estimates of the second-order statics and are, thus, limited to signals with finite variance, the developed distributed particle filtering approach is suitable for the generality of α-stable signals, allowing the proposed algorithm to be used in a multitude of applications. Finally, the so introduced distributed particle filtering approach is validated through a simulation example.
Sayed Pouria Talebi, Danilo P. Mandic
IEEE Signal Process. Lett.2
2017 Complementary Mean Square Analysis of Augmented CLMS for Second-Order Noncircular Gaussian Signals
abstract
A novel physical insight is provided into the behavior and performance of the augmented complex least mean square (ACLMS) algorithm for widely linear adaptive estimation of general second-order noncircular (improper) Gaussian signals, whereby the off-diagonal elements of the covariance and complementary covariance matrices are nonzero. This is achieved through a novel complementary mean square analysis, a counterpart to the standard mean square analysis, and which focuses on the behavior of the complementary second-order statistics of the output error and the augmented weight error vector. We next establish the effect of the degree of input noncircularity on the evolution for these two key parameters that govern the ACLMS. Both transient and steady-state performances are addressed and a stability bound on the step-size for their convergence is established. Simulations in the system identification setting support the analysis.
Yili Xia, Danilo P. Mandic
IEEE Signal Process. Lett.2
2016 Blind source separation and artefact cancellation for single channel bioelectrical signal
abstract
Bioelectrical signal analysis is gaining significant interests from both academics and industries due to its capability for improved diagnosis and therapy of chronic diseases. In practice, different bio-signals, such as EEG, ECG, EOG and EMG, are usually contaminating each other, and the measured signal is the linear combination of them. It is critical to separate them since analysis of one type or several of them separately is of more interest. In the case of multichannel recording, several blind source separation methods are available to extract its original components. However, for single channel scenarios, the problem has yet to be well studied. Therefore in this paper, we explore blind source separation and artefact cancellation for a single channel signal by combining signal decomposition method singular spectrum analysis (SSA) with different blind source separation methods, such as principal component analysis (PCA), maximum noise fraction (MNF), independent component analysis (ICA) and canonical correlation analysis (CCA). We also systematically compare the separation performance by combing different decomposition methods (wavelet transform (WT), ensemble empirical mode decomposition (EEMD) and SSA) with blind source separation methods (PCA, MNF ICA and CCA). The good simulation results have demonstrated the effectiveness and efficiency of the proposed method.
Zhiqiang Zhang 0001, Danilo P. Mandic
BSN3
2016 Modelling stress in public speaking: Evolution of stress levels during conference presentations
abstract
The Electrocardiogram (ECG) collected in real-life scenarios is often noisy and contaminated with motion artefacts. This study proposes a new framework to analyse the heart rate variability (HRV) in mobile scenarios by introducing novel R-peak detection and HRV detrending algorithms. The R-peak detection combines matched filtering and Hilbert transform, while detrending the HRV is performed using empirical mode decomposition with novel physically meaningful stopping criteria. Next, four quantitative metrics-sample entropy, LFhrv, HFhrv and LF/HF ratio - are used to estimate stress levels in two public speaking events: (i) a presentation in front of an audience and (ii) an interactive poster presentation, both at ICASSP 2015. We show that the proposed framework makes it possible to detect distinctive `stress-patterns' in the structural complexity of the HRV, thus verifying the complexity-loss hypothesis in physiological research.
Theerasak Chanwimalueang, Lisa Aufegger, Wilhelm von Rosenberg, Danilo P. Mandic
ICASSP4
2016 Stability analysis of the least-mean-magnitude-phase algorithm
abstract
The least-mean-magnitude-phase (LMMP) algorithm is useful for complex-valued signal processing applications where control of magnitude and/or phase error information is needed to achieve good overall performance. Due to the highly-nonlinear nature of the update terms in the LMMP and related methods, few convergence and stability results exist to guide step size choices for such algorithms. In this paper, we provide a rigorous stability and convergence analysis of the LMMP algorithm using robustness procedures and give sufficient stability conditions on the magnitude and phase step sizes to guarantee contraction-mapping behavior. We also provide an approximate relation for the steady-state MSE as a function of the step size values. Simulations verify the predictive powers of our analytical results and yield useful insights on step size choices in practice.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2016 Novel quaternion matrix factorisations
abstract
The recent introduction of η-Hermitian matrices A = AηHhas opened a new avenue of research in quaternion signal processing. However, the exploitation of this matrix structure has been limited, perhaps due to the lack of joint diagonalisation methodologies of these matrices. As such, we propose novel decompositions of η-Hermitian matrices to address this shortcoming in the literature. As an application, we consider a blind source separation problem in the form of an Alamouti-based communication system. Simulation studies demonstrate the effectiveness of our proposed joint diago-nalisation technique and indicate that our approach is particularly useful when the sources are correlated.
Shirin Enshaeifar, Clive Cheong Took, Saeid Sanei, Danilo P. Mandic
ICASSP4
2016 Quantifying cooperation in choir singing: Respiratory and cardiac synchronisation
abstract
Cooperative tasks require coordinated joint actions among the participants, to the extent that a failure in an individual's action may have catastrophic consequences on the task of the group as a whole. One such activity is choir singing, where highly synchronised performance of the individual singers is a prerequisite to successful performance. The aim of this work is to provide a quantitative measure of the level of cooperation, established through the degrees of synchronisation between singers' physiological responses. To this end, we employ two new measures, the intrinsic phase synchrony and intrinsic coherence, which quantify synchronisation in respiration and heart rate variability (HRV) of: (i) five members of a choir and the conductor during a rehearsal and a real performance, and (ii) five members of the audience attending the performance. Both the proposed techniques successfully reveal degrees of synchronisation of singers' physiological signals which can be used as physically meaningful measures of the level of cooperation.
Apit Hemakom, Valentin Goverdovsky, Lisa Aufegger, Danilo P. Mandic
ICASSP4
2016 Performance advantage of quaternion widely linear estimation: An approximate uncorrelating transform approach
abstract
Widely linear processing has been shown to be superior to the traditional strictly linear processing in quaternion minimum mean square error (MMSE) estimation. However, a quantifiable performance difference between strictly and widely linear processing and the relationship between the performance and quaternion impropriety are still lacking. To this end, we present a proof for the performance advantage of widely linear estimation and relate the performance bounds to signal properties by exploiting the approximate joint diagonalisation of quaternion covariance matrices. In that sense, this work can be seen as a generalisation of complex-valued MMSE estimation, and can thus also be applied to the complex-valued case. Simulations on synthetic signals support the analysis.
Min Xiang, Sithan Kanna, Scott C. Douglas, Danilo P. Mandic
ICASSP4
2016 A least squares enhanced smart DFT technique for frequency estimation of unbalanced three-phase power systems
abstract
The problem of off-nominal frequency estimation in unbalanced three-phase power systems is addressed from the frequency domain perspective. It is first established that the original Smart discrete Fourier transform (SDFT) technique designed for real-valued single-phase voltage can be extended to deal with complex-valued αβ transformed voltage. By observing that the underlying time series relationship among the consecutive DFT fundamental components employed by SDFT technique does not hold when noise or unexpected higher order harmonics are present, resulting in suboptimal estimation performances, the least squares framework is then built upon the underlying relationship among the consecutive DFT fundamental components to minimise the mean square model error. The benefits of the proposed LS-SDFT over the time-domain widely linear least squares (WL-LS) frequency estimator are verified by simulations for unbalanced power system conditions in the presence of noise and higher order harmonic pollution, as well as for real-world measurements.
Yili Xia, Kai Wang 0020, Wenjiang Pei, Zoran Blazic, Danilo P. Mandic
IJCNN5
2016 Discriminating Multiple Emotional States from EEG Using a Data-Adaptive, Multiscale Information-Theoretic Approach
abstract
A multivariate sample entropy metric of signal complexity is applied to EEG data recorded when subjects were viewing four prior-labeled emotion-inducing video clips from a publically available, validated database. Besides emotion category labels, the video clips also came with arousal scores. Our subjects were also asked to provide their own emotion labels. In total 30 subjects with age range 19-70 years participated in our study. Rather than relying on predefined frequency bands, we estimate multivariate sample entropy over multiple data-driven scales using the multivariate empirical mode decomposition (MEMD) technique and show that in this way we can discriminate between five self-reported emotions (p < 0.05). These results could not be obtained by analyzing the relation between arousal scores and video clips, signal complexity and arousal scores, and self-reported emotions and traditional power spectral densities and their hemispheric asymmetries in the theta, alpha, beta, and gamma frequency bands. This shows that multivariate, multiscale sample entropy is a promising technique to discriminate multiple emotional states from EEG recordings.
Yelena Tonoyan, David Looney, Danilo P. Mandic, Marc M. Van Hulle
Int. J. Neural Syst.3
2016 Steady-State Behavior of General Complex-Valued Diffusion LMS Strategies
abstract
A novel methodology to bound the steady-state mean square performance of the diffusion complex least mean square (D-CLMS) and the diffusion widely linear (augmented) CLMS (D-ACLMS) algorithm is proposed. This is achieved by exploiting the almost identical nature of the steady-state filter weights at all nodes. The proposed approach allows for the consideration of the second-order terms in the recursion for the weight error covariance matrix, without compromising the mathematical tractability of the problem. The closed form expressions for the mean square deviation (MSD) and excess mean square error (EMSE) for both the D-CLMS and D-ACLMS allow for the performance of the algorithms to be quantified as a function of the noncircularity of the input data.
Sithan Kanna, Danilo P. Mandic
IEEE Signal Process. Lett.2
2016 Optimization in Quaternion Dynamic Systems: Gradient, Hessian, and Learning Algorithms
abstract
The optimization of real scalar functions of quaternion variables, such as the mean square error or array output power, underpins many practical applications. Solutions typically require the calculation of the gradient and Hessian. However, real functions of quaternion variables are essentially nonanalytic, which are prohibitive to the development of quaternion-valued learning systems. To address this issue, we propose new definitions of quaternion gradient and Hessian, based on the novel generalized Hamilton-real (GHR) calculus, thus making a possible efficient derivation of general optimization algorithms directly in the quaternion field, rather than using the isomorphism with the real domain, as is current practice. In addition, unlike the existing quaternion gradients, the GHR calculus allows for the product and chain rule, and for a one-to-one correspondence of the novel quaternion gradient and Hessian with their real counterparts. Properties of the quaternion gradient and Hessian relevant to numerical applications are also introduced, opening a new avenue of research in quaternion optimization and greatly simplified the derivations of learning algorithms. The proposed GHR calculus is shown to yield the same generic algorithm forms as the corresponding real- and complex-valued algorithms. Advantages of the proposed framework are illuminated over illustrative simulations in quaternion signal processing and neural networks.
Dongpo Xu, Yili Xia, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2016 Is a Complex-Valued Stepsize Advantageous in Complex-Valued Gradient Learning Algorithms?
abstract
Complex gradient methods have been widely used in learning theory, and typically aim to optimize real-valued functions of complex variables. The stepsize of complex gradient learning methods (CGLMs) is a positive number, and little is known about how a complex stepsize would affect the learning process. To this end, we undertake a comprehensive analysis of CGLMs with a complex stepsize, including the search space, convergence properties, and the dynamics near critical points. Furthermore, several adaptive stepsizes are derived by extending the Barzilai-Borwein method to the complex domain, in order to show that the complex stepsize is superior to the corresponding real one in approximating the information in the Hessian. A numerical example is presented to support the analysis.
Huisheng Zhang, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.2
2016 Group Component Analysis for Multiblock Data: Common and Individual Feature Extraction
abstract
Real-world data are often acquired as a collection of matrices rather than as a single matrix. Such multiblock data are naturally linked and typically share some common features while at the same time exhibiting their own individual features, reflecting the underlying data generation mechanisms. To exploit the linked nature of data, we propose a new framework for common and individual feature extraction (CIFE) which identifies and separates the common and individual features from the multiblock data. Two efficient algorithms termed common orthogonal basis extraction (COBE) are proposed to extract common basis is shared by all data, independent on whether the number of common components is known beforehand. Feature extraction is then performed on the common and individual subspaces separately, by incorporating dimensionality reduction and blind source separation techniques. Comprehensive experimental results on both the synthetic and real-world data demonstrate significant advantages of the proposed CIFE method in comparison with the state-of-the-art.
Guoxu Zhou, Andrzej Cichocki, Yu Zhang 0009, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.4
2015 Nonuniformly sampled trivariate empirical mode decomposition
abstract
Multichannel data-driven time-frequency algorithms, such as the multivariate empirical mode decomposition (MEMD), have emerged as important tools in the analysis of inter-channel dependencies that arise in multivariate data. Such methods employ uniform projection schemes on hyperspheres in order to estimate the local mean, thus requiring dense but underutilised sampling when processing unbalanced data channels. To this end, we propose a nonuniform projection scheme that adapts to the second order statistics of trivariate data; this provides the estimation of the local mean in the case of power imbalances and correlations between the channels. The algorithm is particularly useful for generating a low number of direction vectors within MEMD. Its performance is illustrated on synthetic and real-world data.
Apit Hemakom, Alireza Ahrabian, David Looney, Naveed ur Rehman, Danilo P. Mandic
ICASSP5
2015 Mean square analysis of the CLMS and ACLMS for non-circular signals: The approximate uncorrelating transform approach
abstract
Current approaches to the mean-square analyses of the complex-least-mean-square (CLMS) and augmented CLMS (ACLMS) algorithms can be challenging due to the difficulty in diagonalising the augmented covariance matrix. By employing the recently introduced approximate uncorrelating transform (AUT), which diagonalizes the covariance and pseudocovariance matrices with a single singular value decomposition (SVD), we derive closed form expressions for both transient and steady-state mean square stability for the CLMS and ACLMS. Relationships between the degree of circularity of the input signal and the bound on the step-sizes of the CLMS and ACLMS are also established. We also show that for both CLMS and ACLMS, the steady-state misadjustment increases with the degree of non-circularity of the input signal. Simulations in the context of frequency estimation in power grid support the analyses.
Danilo P. Mandic, Sithan Kanna, Scott C. Douglas
ICASSP1
2015 Vital signs from inside a helmet: A multichannel face-lead study
abstract
It is essential to measure physiological parameters such as heart rate variability and respiratory rate of drivers to evaluate their performance. The results from this measurement can be used to assess the state of body and mind, for instance concentration and stress. However, current systems only work in controlled environments, or sensors obstruct and interfere with operations of the driver. In this study, a face-lead ECG is placed inside a helmet to enhance comfort and convenience in racing scenarios. Multiple electrodes were attached to facial locations, which exhibit good contact with a helmet, and bipolar configurations were examined between the left and right side of the subject's face. Standard and data-driven filtering algorithms were employed to improve the extraction of R peaks from the ECG data. The so-extracted R peaks were subsequently used to estimate heart activity and respiration effort, and the results were compared with standard recording protocols. It is shown that ECG recordings obtained from locations on the lower jaw match closely with conventional recording paradigms (limb-lead ECG), highlighting the potential of vital sign monitoring from within a racing helmet.
Wilhelm von Rosenberg, Theerasak Chanwimalueang, David Looney, Danilo P. Mandic
ICASSP4
2015 A quaternion frequency estimator for three-phase power systems
abstract
Motivated by the need for accurate frequency estimation in power systems, a novel algorithm for estimating the fundamental frequency of both balanced and unbalanced three-phase power systems, which is robust to noise and distortions, is developed. This is achieved through the use of quaternions in order to provide a unified framework for joint modeling of voltage measurements from all the phases of a three-phase system. Next, the recently introduced Hℝ-calculus is employed to derive a state space estimator based on the quaternion extended Kalman filter (QEKF). The proposed algorithm is validated over a variety of case studies using both synthetic and real-world data.
Sayed Pouria Talebi, Danilo P. Mandic
ICASSP2
2015 The widely linear quaternion recursive total least squares
abstract
A widely linear quaternion recursive total least squares (WL-QRTLS) algorithm is introduced for the processing of Q-improper processes contaminated by noise. The total least squares for quaternions (QTLS) is a generalisation of the real-valued total least squares and is introduced rigorously, starting from the existence condition for low-rank approximation of quaternion matrices. Then, a quaternion Rayleigh quotient (QRQ) is defined to establish the link between the QTLS solution and the minimisation of the QRQ. Finally, the rank-one update formula is employed to allow for fast iterative solution based on the QRQ. Through simulations, the WL-QRTLS was shown to exhibit superior performance, under perturbations on both input and output signals, to other adaptive filtering of the same class - the widely linear quaternion least mean squares (WL-QLMS) and the widely linear quaternion recursive least squares (WL-QRLS). The experiments on both synthetic and real-world $\mathbb{Q}$-improper processes supported the analysis.
Thiannithi Thanthawaritthisai, Felipe A. Tobar, Anthony G. Constantinides, Danilo P. Mandic
ICASSP4
2015 Common components analysis via linked blind source separation
abstract
Very often data we encounter in practice is a collection of matrices rather than a single matrix. These multi-block data often share some common features, due to the background in which they are measured. In this study we propose a new concept of linked blind source separation (BSS) that aims at discovering and extracting unique and physically meaningful common components from multi-block data, which also contain strong individual components. The validity and potential of the proposed method is justified by simulations.
Guoxu Zhou, Andrzej Cichocki, Danilo P. Mandic
ICASSP3
2015 A non-linear state space frequency estimator for three-phase power systems
abstract
Frequency estimation in three-phase power systems is considered from a state space point of view, and a robust and fast converging algorithm for estimating the fundamental frequency of three-phase power systems is introduced. This is achieved by exploiting the Clarke transform to incorporate the information from all the phases and then designing a widely linear state space estimator that can accurately estimate the fundamental frequency of both balanced and unbalanced three-phase power systems. The framework is then expanded to modify the state space model in order to account for the presence of harmonics in the system. The performance of the developed algorithm is validated through simulations on both synthetic data and real-world data recordings, where it is shown that the developed algorithm outperforms standard linear and the recently introduced widely liner frequency estimators.
Sayed Pouria Talebi, Sithan Kanna, Danilo P. Mandic
IJCNN3
2015 Convergence analysis of an augmented algorithm for fully complex-valued neural networks
Dongpo Xu, Huisheng Zhang, Danilo P. Mandic
Neural Networks3
2015 Synchrosqueezing-based time-frequency analysis of multivariate data
Alireza Ahrabian, David Looney, Ljubisa Stankovic, Danilo P. Mandic
Signal Process.4
2015 Selective Time-Frequency Reassignment Based on Synchrosqueezing
abstract
Reassignment methods seek to sharpen the time-frequency representation of conventional time-frequency algorithms, such as the continuous wavelet transform (CWT). However, such methods aim to localize both noise components and signal components of interest, which makes the discrimination between such components for low SNR signals a difficult task. Inspired by the recovery of modes (RCM) algorithm, we propose a selective time-frequency reassignment procedure that attempts to identify and localize oscillatory components of interest for the continuous wavelet transform (CWT), where the reassignment is carried out for selective localization. The performance of the proposed method is illustrated on both synthetic and real world data.
Alireza Ahrabian, Danilo P. Mandic
IEEE Signal Process. Lett.2
2015 Design of Positive-Definite Quaternion Kernels
abstract
Quaternion reproducing kernel Hilbert spaces (QRKHS) have been proposed recently and provide a high-dimensional feature space (alternative to the real-valued multikernel approach) for general kernel-learning applications. The current challenge within quaternion-kernel learning is the lack of general quaternion-valued kernels, which are necessary to exploit the full advantages of the QRKHS theory in real-world problems. This letter proposes a novel way to design quaternion-valued kernels, this is achieved by transforming three complex kernels into quaternion ones and then combining their real and imaginary parts. Building on this general construction, our emphasis is on a new quaternion kernel of polynomial features, which is assessed in the prediction of bodysensor networks applications.
Felipe A. Tobar, Danilo P. Mandic
IEEE Signal Process. Lett.2
2015 A Multivariate Empirical Mode DecompositionBased Approach to Pansharpening
abstract
We propose a novel class of schemes for the pansharpening of multispectral (MS) images using a multivariate empirical mode decomposition (MEMD) algorithm. MEMD is an extension of the empirical mode decomposition (EMD) algorithm, which enables the decomposition of multivariate data into its intrinsic oscillatory scales. The ability of MEMD to process multichannel data directly by performing data-driven, local, and multiscale analysis makes it a perfect match for pansharpening applications, a task for which standard univariate EMD is ill-equipped due to the nonuniqueness, mode-mixing, and mode-misalignment issues. We show that MEMD overcomes the limitations of standard EMD and yields improved spatial and spectral performance in the context of pansharpening of MS images. The potential of the proposed schemes is further demonstrated through comparative analysis against a number of standard pansharpening algorithms on both simulated Pleiades and real-world IKONOS data sets.
Syed Muhammad Umer Abdullah, Naveed ur Rehman, Muhammad Murtaza Khan, Danilo P. Mandic
IEEE Trans. Geosci. Remote. Sens.4
2015 Maintaining the Integrity of Sources in Complex Learning Systems: Intraference and the Correlation Preserving Transform
abstract
The correlation preserving transform (CPT) is introduced to perform bivariate component analysis via decorrelating matrix decompositions, while at the same time preserving the integrity of original bivariate sources. Specifically, unlike existing bivariate uncorrelating matrix decomposition techniques, CPT is designed to preserve both the order of the data channels within every bivariate source and their mutual correlation properties. We introduce the notion of intraference to quantify the effects of interchannel mixing artifacts within recovered bivariate sources, and show that the integrity of separated sources is compromised when not accounting for the intrinsic correlations within bivariate sources, as is the case with current bivariate matrix decompositions. The CPT is based on augmented complex statistics and involves finding the correct conjugate eigenvectors associated with the pseudocovariance matrix, making it possible to maintain the physical meaning of the separated sources. The benefits of CPT are illustrated in the source separation and clustering scenarios, for both synthetic and real-world data.
Clive Cheong Took, Scott C. Douglas, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2015 Quaternion-Valued Echo State Networks
abstract
Quaternion-valued echo state networks (QESNs) are introduced to cater for 3-D and 4-D processes, such as those observed in the context of renewable energy (3-D wind modeling) and human centered computing (3-D inertial body sensors). The introduction of QESNs is made possible by the recent emergence of quaternion nonlinear activation functions with local analytic properties, required by nonlinear gradient descent training algorithms. To make QENSs second-order optimal for the generality of quaternion signals (both circular and noncircular), we employ augmented quaternion statistics to introduce widely linear QESNs. To that end, the standard widely linear model is modified so as to suit the properties of dynamical reservoir, typically realized by recurrent neural networks. This allows for a full exploitation of second-order information in the data, contained both in the covariance and pseudocovariances, and a rigorous account of second-order noncircularity (improperness), and the corresponding power mismatch and coupling between the data components. Simulations in the prediction setting on both benchmark circular and noncircular signals and on noncircular real-world 3-D body motion data support the analysis.
Yili Xia, Cyrus Jahanchahi, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2015 Performance Bounds of Quaternion Estimators
abstract
The quaternion widely linear (WL) estimator has been recently introduced for optimal second-order modeling of the generality of quaternion data, both second-order circular (proper) and second-order noncircular (improper). Experimental evidence exists of its performance advantage over the conventional strictly linear (SL) as well as the semi-WL (SWL) estimators for improper data. However, rigorous theoretical and practical performance bounds are still missing in the literature, yet this is crucial for the development of quaternion valued learning systems for 3-D and 4-D data. To this end, based on the orthogonality principle, we introduce a rigorous closed-form solution to quantify the degree of performance benefits, in terms of the mean square error, obtained when using the WL models. The cases when the optimal WL estimation can simplify into the SWL or the SL estimation are also discussed.
Yili Xia, Cyrus Jahanchahi, Tohru Nitta, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.4
2014 Estimation of financial indices volatility using a model with time-varying parameters
abstract
A class of stochastic volatility models (SVMs) with time-varying parameters is presented for online volatility estimation in nonstationary environments. This is achieved by modelling both the volatility and model parameters as states of a hidden Markov model (HMM), thus allowing for the use of particle filters to estimate the resulting posterior densities. The proposed models, based on the logarithmic SVM and the unobserved GARCH model, are evaluated for the estimation of the volatility of the NASDAQ-C and the Chilean IGPA financial indices between June 2007 and January 2010, where the late-2000s financial crisis is included. Simulations show that the proposed time-varying models are well suited for online volatility estimation as (i) they achieve an accuracy comparable to those of offline (batch) algorithms, and (ii) their parameters can be used to identify market changes.
Felipe A. Tobar, Marcos E. Orchard, Danilo P. Mandic, Anthony G. Constantinides
CIFEr3
2014 Estimation of phase synchrony using the synchrosqueezing transform
abstract
Phase synchronization has emerged as an important concept in quantifying interactions between dynamical systems. In this work a robust estimate of the phase synchrony between bivariate signals is presented. This is achieved by extending the recently introduced synchrosqueezing transform (SST), a method that belongs to the class of reassignment techniques that generates highly localized time-frequency representations, so as to cater for bivariate data. The proposed method is shown to generate accurate estimates of phase synchrony on both synthetic and real world signals.
Alireza Ahrabian, Danilo P. Mandic
ICASSP2
2014 Autoconvolution and panorama: Augmenting second-order signal analysis
abstract
The autocorrelation does not differentiate between deterministic and stochastic signals, as phase information is not maintained. This paper introduces the autoconvolution for both deterministic and stochastic signals. The autoconvolution with the autocorrelation provides a second-order description that discriminates between deterministic and stochastic signals - even those with identical power spectra. We also introduce the panorama as the Fourier transform of the autoconvolution. The power spectrum and panorama admit a two-dimensional spectral representation that has unique and powerful properties, such as detecting deterministic sinusoidal components in correlated stochastic noise without knowledge of the sinusoidal frequencies or amplitudes. Additional extensions are indicated.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2014 Subspace denoising of EEG artefacts via multivariate EMD
abstract
The components obtained using the time-frequency algorithm empirical mode decomposition (EMD) enable unique advantages in the context of noise removal. In this paper, recent EMD-based de-noising methods are reviewed and similarities with a more conventional class of techniques - subspace denoising - are illustrated. Standard subspace approaches which are based on the factorisation of covariance matrices are unsuitable for nonstationary data. By comparison, EMD facilitates a signal representation which enables denoising using short spatio/temporal windows. It is highlighted how the EMD property of local orthogonality can be extended via multivariate operations, and a denoising scheme is proposed and compared with standard methods in electroencephalogram (EEG) artefact-removal using a novel multimodal sensor.
David Looney, Valentin Goverdovsky, Preben Kidmose, Danilo P. Mandic
ICASSP4
2014 A Quaternion Least Mean Phase adaptive estimator
abstract
A Quaternion Least Mean Phase (QLMP) algorithm is introduced for phase-only adaptive filtering of quaternion-valued signals. This is achieved by first defining the notion of phase in the quaternion domain based on the exponential representation of quaternion random variables, and then using the HR-calculus to derive a steepest-decent weight update. The advantages of this adaptive algorithm are illustrated in a bearings only tracking scenario, where the QLMP is shown to outperform the amplitude-phase based quaternion Least Mean Square (QLMS).
Sayed Pouria Talebi, Dongpo Xu, Anthony Kuh, Danilo P. Mandic
ICASSP4
2014 A particle filtering based kernel HMM predictor
abstract
A novel kernel algorithm is proposed for nonlinear prediction whereby the signal is modelled as a state of a hidden Markov model (HMM). The transition function of the HMM is approximated using kernels, whose weights are also part of the state of the system and are learnt in an unsupervised fashion by a sample importance resampling (SIR) particle filter. The SIR proposal density is designed so as to maintain a diverse population of particles, thus avoiding particle degeneracy arising from inaccuracies of early model estimates. The kernel HMM algorithm is further equipped with a sparsification criterion based on approximate linear dependence and its performance is evaluated against the KNLMS and KRLS algorithms for the prediction of synthetic signals and real world point-of-gaze data.
Felipe A. Tobar, Danilo P. Mandic
ICASSP2
2014 The HC calculus, quaternion derivatives and caylay-hamilton form of quaternion adaptive filters and learning systems
abstract
We introduce a novel and unifying framework for the calculation of gradients of both quaternion holomorphic functions and nonholomorphic real functions of quaternion variables. This is achieved by considering the isomorphism between the quaternion domain H and the bivariate complex domain C×C, and by exploiting complex calculus to simplify the quaternion gradient calculation. The validation of the proposed HC calculus is performed against the existing HR calculus, and its convenience is illustrated in the context of gradient-based quaternion optimisation as well as in adaptive learning systems. Quaternion adaptive filtering algorithms and a dynamical perceptron update are next derived based on the bivariate complex representation of quaternions and the HC calculus. Simulations on both synthetic and real-world multidimensional signals support the analysis.
Yili Xia, Cyrus Jahanchahi, Dongpo Xu, Danilo P. Mandic
IJCNN4
2014 Complex dual channel estimation: Cost effective widely linear adaptive filtering
Cyrus Jahanchahi, Sithan Kanna, Danilo P. Mandic
Signal Process.3
2014 Dynamically-Sampled Bivariate Empirical Mode Decomposition
abstract
A novel scheme for selecting projection vectors in bivariate empirical mode decomposition (BEMD) is proposed in order to enable accurate signal decomposition at lower computational complexity. Unlike existing algorithms which use a static uniform scheme for the distribution of projection vectors, the proposed scheme examines local curvature in multidimensional spaces to produce a data-adaptive set of direction vectors for taking signal projections. This is achieved by aligning the density of projection vectors according to the empirical distributions of angles where the signal exhibits highest local dynamics. We show that the proposed methodology outperforms the existing schemes for a small number of signal projections. The proposed algorithm is verified via illustrative simulations demonstrating accurate local mean estimation and mode extraction.
Naveed ur Rehman, Muhammad Waqas Safdar, Danilo P. Mandic
IEEE Signal Process. Lett.4
2014 Quaternion Reproducing Kernel Hilbert Spaces: Existence and Uniqueness Conditions
abstract
The existence and uniqueness conditions of quaternion reproducing kernel Hilbert spaces (QRKHS) are established in order to provide a mathematical foundation for the development of quaternion-valued kernel learning algorithms. This is achieved through a rigorous account of left quaternion Hilbert spaces, which makes it possible to generalise standard RKHS to quaternion RKHS. Quaternion versions of the Riesz representation and Moore-Aronszajn theorems are next introduced, thus underpinning kernel estimation algorithms operating on quaternion-valued feature spaces. The difference between the proposed quaternion kernel concept and the existing real and vector approaches is also established in terms of both theoretical advantages and computational complexity. The enhanced estimation ability of the so-introduced quaternion-valued kernels over their real- and vector-valued counterparts is validated through kernel ridge regression applications. Simulations on real world 3D inertial body sensor data and nonlinear channel equalisation using novel quaternion cubic and Gaussian kernels support the approach.
Felipe A. Tobar, Danilo P. Mandic
IEEE Trans. Inf. Theory2
2014 Guest Editorial Special Issue on Complex- and Hypercomplex-Valued Neural Networks
abstract
The fifteen papers in this special issue focus on complex and hyper-complex neural network applications. Complex-valued neural networks (CVNNs) exhibit very desirable characteristics in their learning, self-organizing, and processing dynamics, which makes them attractive for applications in various areas in science and technology. For example, they are perfectly suited to deal with complex amplitude, composed of amplitude and phase, which is one of the core concepts in physical systems dealing with electromagnetic, light, sonic/ultrasonic, and quantum waves.
Akira Hirose 0001, Igor N. Aizenberg, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2014 A Class of Quaternion Kalman Filters
abstract
The existing Kalman filters for quaternion-valued signals do not operate fully in the quaternion domain, and are combined with the real Kalman filter to enable the tracking in 3-D spaces. Using the recently introduced HR-calculus, we develop the fully quaternion-valued Kalman filter (QKF) and quaternion-extended Kalman filter (QEKF), allowing for the tracking of 3-D and 4-D signals directly in the quaternion domain. To consider the second-order noncircularity of signals, we employ the recently developed augmented quaternion statistics to derive the widely linear QKF (WL-QKF) and widely linear QEKF (WL-QEKF). To reduce computational requirements of the widely linear algorithms, their efficient implementation are proposed and it is shown that the quaternion widely linear model can be simplified when processing 3-D data, further reducing the computational requirements. Simulations using both synthetic and real-world circular and noncircular signals illustrate the advantages offered by widely linear over strictly linear quaternion Kalman filters.
Cyrus Jahanchahi, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.2
2014 Multikernel Least Mean Square Algorithm
abstract
The multikernel least-mean-square algorithm is introduced for adaptive estimation of vector-valued nonlinear and nonstationary signals. This is achieved by mapping the multivariate input data to a Hilbert space of time-varying vector-valued functions, whose inner products (kernels) are combined in an online fashion. The proposed algorithm is equipped with novel adaptive sparsification criteria ensuring a finite dictionary, and is computationally efficient and suitable for nonstationary environments. We also show the ability of the proposed vector-valued reproducing kernel Hilbert space to serve as a feature space for the class of multikernel least-squares algorithms. The benefits of adaptive multikernel (MK) estimation algorithms are illuminated in the nonlinear multivariate adaptive prediction setting. Simulations on nonlinear inertial body sensor signals and nonstationary real-world wind signals of low, medium, and high dynamic regimes support the approach.
Felipe A. Tobar, Sun-Yuan Kung, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.3
2014 Adaptive Convex Combination Approach for the Identification of Improper Quaternion Processes
abstract
Data-adaptive optimal modeling and identification of real-world vector sensor data is provided by combining the fractional tap-length (FT) approach with model order selection in the quaternion domain. To account rigorously for the generality of such processes, both second-order circular (proper) and noncircular (improper), the proposed approach in this paper combines the FT length optimization with both the strictly linear quaternion least mean square (QLMS) and widely linear QLMS (WL-QLMS). A collaborative approach based on QLMS and WL-QLMS is shown to both identify the type of processes (proper or improper) and to track their optimal parameters in real time. Analysis shows that monitoring the evolution of the convex mixing parameter within the collaborative approach allows us to track the improperness in real time. Further insight into the properties of those algorithms is provided by establishing a relationship between the steady-state error and optimal model order. The approach is supported by simulations on model order selection and identification of both strictly linear and widely linear quaternion-valued systems, such as those routinely used in renewable energy (wind) and human-centered computing (biomechanics).
Bukhari Che Ujang, Cyrus Jahanchahi, Clive Cheong Took, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.4
2013 Dynamical complexity analysis of multivariate financial data
abstract
Characterization of joint dynamics of multivariate financial time series calls for the analysis based on joint intrinsic temporal and information-theoretic scales. Yet a rigorous account of dynamical complexity of such time series is hampered by the univariate natures and mathematical artefacts associated with the existing methods. To that end, we employ multi-variate multiscale entropy (MMSE) in order to associate multivariate complexity with long-range correlations, direct and indirect couplings, and synchronies among the data channels. Simulations on major stock indices support the approach.
Wenjun Er, Danilo P. Mandic
ICASSP2
2013 The quaternion kernel least squares
abstract
The quaternion kernel least squares algorithm (QKLS) is introduced as a generic kernel framework for the estimation of multivariate quaternion valued signals. This is achieved based on the concepts of quaternion inner product and quaternion positive definiteness, allowing us to define quaternion kernel regression. Next, the least squares solution is derived using the recently introduced Hℝ calculus. We also show that QKLS is a generic extension of standard kernel least squares, and their equivalence is established for real valued kernels. The superiority of the quaternion-valued linear kernel with respect to its real-valued counterpart is illustrated for both synthetic and real-world prediction applications, in terms of accuracy and robustness to overfitting.standard kernel least squares,quaternion-valued linear kernelreal-world prediction applications,real-world 3D inertial body sensor signals.synthetic autoregressive processes
Felipe A. Tobar, Danilo P. Mandic
ICASSP2
2013 On intraference and its implications in complex-valued signal processing
abstract
We introduce the concept of intra-ference in order to quantify the degree to which the integrity of bivariate (or complex) sources is preserved in applications based on matrix decompositions of bivariate data. This is achieved by examining the pseudocovariance matrix of noncircular complex sources, and by recognising that the pseudocovariance is intrinsically complex valued. We illuminate how the existing decompositions such as the strong uncorrelating transform (SUT) not only decorrelate the bivariate sources from one another, but also decorrelate and scatter the data channels within each bivariate source, thus violating source integrity. Examples showing that the intra-ference arises due to the phase ambiguity in the existing matrix decompositions support the approach.
Clive Cheong Took, Scott C. Douglas, Danilo P. Mandic
ICASSP3
2013 Design of oversampled generalised discrete Fourier transform filter banks for application to subband-based blind source separation
abstract
A novel design of oversampled generalised discrete Fourier transform filter banks is proposed, with application to subband‐based convolutive blind source separation (BSS), where either instantaneous BSS algorithms or joint BSS algorithms can be applied. Conventional filter banks design is usually focused on elimination of the overall aliasing error and the perfect reconstruction (PR) condition, which are required by traditional subband adaptive filtering applications. However, because of the unknown scaling factor, the traditional PR condition is not necessary in the context of subband BSS and can be relaxed in the design. Owing to the increased degrees of design freedom, the authors can introduce an additional cost function to enhance the mutual information between adjacent subband signals. Together with a reduced subband aliasing level, it leads to an improved subband permutation alignment result for instantaneous BSS and an overall better performance for the joint BSS.
Wei Liu 0001, Danilo P. Mandic
IET Signal Process.3
2013 Higher Order Partial Least Squares (HOPLS): A Generalized Multilinear Regression Method
abstract
A new generalized multilinear regression model, termed the higher order partial least squares (HOPLS), is introduced with the aim to predict a tensor (multiway array) Y from a tensor X through projecting the data onto the latent space and performing regression on the corresponding latent variables. HOPLS differs substantially from other regression models in that it explains the data by a sum of orthogonal Tucker tensors, while the number of orthogonal loadings serves as a parameter to control model complexity and prevent overfitting. The low-dimensional latent space is optimized sequentially via a deflation operation, yielding the best joint subspace approximation for both X and Y. Instead of decomposing X and Y individually, higher order singular value decomposition on a newly defined generalized cross-covariance tensor is employed to optimize the orthogonal loadings. A systematic comparison on both synthetic data and real-world decoding of 3D movement trajectories from electrocorticogram signals demonstrate the advantages of HOPLS over the existing methods in terms of better predictive ability, suitability to handle small sample sizes, and robustness to noise.
Qibin Zhao, Cesar F. Caiafa, Danilo P. Mandic, Zenas C. Chao, Yasuo Nagasaka, Naotaka Fujii, Liqing Zhang 0001, Andrzej Cichocki
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 A class of quaternion valued affine projection algorithms
Cyrus Jahanchahi, Clive Cheong Took, Danilo P. Mandic
Signal Process.3
2013 Prediction of wide-sense stationary quaternion random signals
Jesús Navarro-Moreno, Rosa M. Fernández-Alcalá, Clive Cheong Took, Danilo P. Mandic
Signal Process.4
2013 Blind Separation of Dependent Sources With a Bounded Component Analysis Deflationary Algorithm
abstract
The problem of blind source separation of complex-valued sources from a linear mixture is addressed. We propose a deflationary algorithm for the sequential recovery of a set of communication signals, where each source is extracted by performing a Bounded Component Analysis of the linear mixture. The contribution of each recovered source to the observations is removed by minimizing its convex perimeter, without using second-order statistics. This implies to run a gradient descent algorithm several times. In order to accelerate the convergence, we have derived a fast step size that exploits the second-order information of the cost function by means of the augmented Hessian matrix. Computer simulations show that the proposed method is able to blindly separate even dependent sources, as long as they satisfy the BCA separability conditions. Also, the speed of convergence of this novel step size is compared with other classical approaches.
Pablo Aguilera, Sergio Cruces, Iván Durán-Díaz, Auxiliadora Sarmiento, Danilo P. Mandic
IEEE Signal Process. Lett.5
2013 Bivariate Empirical Mode Decomposition for Unbalanced Real-World Signals
abstract
The bivariate empirical mode decomposition (BEMD) algorithm employs uniform sampling on a circle to perform projections in multiple directions, in order to calculate the local mean of a bivariate signal. However, this approach is adequate only for equal powers in both the data channels within a bivariate signal, and results in suboptimal performance for data channels exhibiting power imbalance, a typical case in practice. To that end, we exploit second-order bivariate statistical properties to introduce a nonuniform sampling scheme for data adaptive selection of the projection directions. In this way, the resulting nonuniformly sampled BEMD (NS-BEMD) algorithm provides a more accurate time-frequency representation of bivariate data than standard BEMD, for the same number of projections. The advantages of the proposed approach are demonstrated in case studies on BEMD for correlated data channels, selection of optimal noise power in noise-assisted BEMD, and for speed estimation using Doppler radar.
Alireza Ahrabian, Naveed ur Rehman, Danilo P. Mandic
IEEE Signal Process. Lett.3
2012 Multivariate entropy analysis with data-driven scales
abstract
A data-adaptive algorithm for the entropy-based analysis of structural regularities (complexity) in multivariate signals is proposed. This is achieved by combining multivariate sample entropy with a multivariate extension of empirical mode decomposition, both data-driven multiscale techniques. The proposed analysis across data-adaptive scales makes the approach robust to nonstationarity, a critical issue with information theoretic measures. Simulations on synthetic and real-world physiological data support the approach and validate the hypothesis of increased complexity for unconstrained as compared to constrained (due to e.g. ageing or illness) biological systems.
Mosabber Uddin Ahmed, Naveed ur Rehman, David Looney, Tomasz M. Rutkowski, Preben Kidmose, Danilo P. Mandic
ICASSP6
2012 A family of least-squares magnitude phase algorithms
abstract
This paper presents a family of least-squares algorithms for adaptive signal processing of complex-valued signals. The algorithms employ a composite cost function that allows magnitude and phase errors to be weighted differently in the parameter estimation depending on their importance, providing an opportunity for enhanced estimation performance over standard least-squares methods. We also describe a procedure for automatically adjusting this weighting based on the estimation errors themselves. Simulations show the excellent behavior of the algorithms in time-varying signal conditions.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2012 On gradient calculation in quaternion adaptive filtering
abstract
A novel way to calculate the gradient of real functions of quaternion variables, typical cost functions in quaternion signal processing, is proposed. This is achieved by revisiting quaternion involutions and by simplifying the existing HR derivatives. This has allowed us to express the class of quaternion least mean square (QLMS) algorithms in a more compact form while keeping the same generic form of LMS. Simulations in the prediction setting support the approach.
Cyrus Jahanchahi, Clive Cheong Took, Danilo P. Mandic
ICASSP3
2012 On quaternion analyticity: Enabling quaternion-valued nonlinear adaptive filtering
abstract
The strict Cauchy-Riemann-Fueter (CRF) analyticity conditions establish that only linear quaternion-valued functions are analytic, prohibiting the development of quaternion-valued nonlinear adaptive filters for the recurrent neural network architecture (RNN). In this work, the requirement of local analyticity in gradient based learning is exercised and proposes to use the local analyticity condition (LAC) to introduce quaternion-valued nonlinear feedback adaptive filters. The introduced class of algorithms make full use of quaternion algebra and provide generic extensions of the corresponding real and complex solutions. Simulations in the prediction setting support the analysis presented.
Bukhari Che Ujang, Clive Cheong Took, Danilo P. Mandic
ICASSP3
2012 Kalman filtering for widely linear complex and quaternion valued bearings only tracking
abstract
Bearings only target tracking is concerned with estimating the trajectory of an object from noise-corrupted bearing (phase) measurements. Traditionally this problem has been formulated as real valued for the Cartesian coordinate system or modified polar coordinate system. In this study, the authors introduce the bearings only tracking problem for the complex and quaternion domains to take advantage of the natural representation offered by these domains, for multivariate real signals, as well as the greater insights provided into the dynamics of tracking. Moreover, the authors introduce the augmented complex and quaternion extended Kalman filters for the modelling of second-order non-circular complex and quaternion valued signals, for which a widely linear model is shown to be more suitable than a strictly linear model.
Dahir H. Dini, Cyrus Jahanchahi, Danilo P. Mandic
IET Signal Process.3
2012 Reducing permutation error in subband-based convolutive blind separation
abstract
Subband-based blind source separation has great potential in solving the complicated convolutive mixing problem. However, its performance is largely affected by the permutation ambiguity problem during the synthesis stage. Researchers have suggested methods to correct the permutation by exploiting the correlation information between adjacent frequencies/subbands. An improved solution to this permutation problem is proposed based on a novel filter banks design method, which is based on a model that includes inter-subband correlation as part of the optimisation criterion. Simulation results show that a better subband permutation alignment result has been achieved, leading to improved separation performance.
Wei Liu 0001, Danilo P. Mandic
IET Signal Process.3
2012 Modelling of brain consciousness based on collaborative adaptive filters
Ling Li 0010, Yili Xia, Beth Jelfs, Jianting Cao, Danilo P. Mandic
Neurocomputing5
2012 An adaptive approach for the identification of improper complex signals
Beth Jelfs, Danilo P. Mandic, Scott C. Douglas
Signal Process.2
2012 Multivariate Multiscale Entropy Analysis
abstract
Multivariate physical and biological recordings are common and their simultaneous analysis is a prerequisite for the understanding of the complexity of underlying signal generating mechanisms. Traditional entropy measures are maximized for random processes and fail to quantify inherent long-range dependencies in real world data, a key feature of complex systems. The recently introduced multiscale entropy (MSE) is a univariate method capable of detecting intrinsic correlations and has been used to measure complexity of single channel physiological signals. To generalize this method for multichannel data, we first introduce multivariate sample entropy (MSampEn) and evaluate it over multiple time scales to perform the multivariate multiscale entropy (MMSE) analysis. This makes it possible to assess structural complexity of multivariate physical or physiological systems, together with more degrees of freedom and enhanced rigor in the analysis. Simulations on both multivariate synthetic data and real world postural sway analysis support the approach.
Mosabber Uddin Ahmed, Danilo P. Mandic
IEEE Signal Process. Lett.2
2012 Class of Widely Linear Complex Kalman Filters
abstract
Recently, a class of widely linear (augmented) complex-valued Kalman filters (KFs), that make use of augmented complex statistics, have been proposed for sequential state space estimation of the generality of complex signals. This was achieved in the context of neural network training, and has allowed for a unified treatment of both second-order circular and noncircular signals, that is, both those with rotation invariant and rotation-dependent distributions. In this paper, we revisit the augmented complex KF, augmented complex extended KF, and augmented complex unscented KF in a more general context, and analyze their performances for different degrees of noncircularity of input and the state and measurement noises. For rigor, a theoretical bound for the performance advantage of widely linear KFs over their strictly linear counterparts is provided. The analysis also addresses the duality with bivariate real-valued KFs, together with several issues of implementation. Simulations using both synthetic and real world proper and improper signals support the analysis.
Dahir H. Dini, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.2
2011 The least-mean-magnitude-phase algorithm with applications to communications systems
abstract
This paper presents the least-mean-magnitude-phase (LMMP) algorithm for adaptive signal processing of complex-valued signals. The algorithm employs a decomposition of the mean squared error cost and allows different normalized step sizes to be selected for the magnitude and phase errors. Simulations show that the algorithm is useful when either the amplitude or the phase relationship between the input and desired response signals can be accurately estimated.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2011 Blind extraction of improper quaternion sources
abstract
Blind extraction of quaternion-valued latent sources is addressed based on their local temporal properties. The extraction criterion is based on the minimum mean square widely linear prediction error, thus allowing for the extraction of both proper and improper quaternion sources. The use of the widely linear adaptive predictor is justified by the relationship between the mean square prediction error and the crosscorrelation and cross-pseudocorrelations of the source signals. Simulations on benchmark improper quaternion sources together with a real-world example of EEG artifact removal illustrate the usefulness of the proposed methodology. © 2011 IEEE.
Soroush Javidi, Clive Cheong Took, Cyrus Jahanchahi, Nicolas Le Bihan, Danilo P. Mandic
ICASSP5
2011 Augmented complex matrix factorisation
abstract
A novel framework for the factorisation of complex-valued data is derived using recent developments in complex statistics. Unlike existing factorisation tools the algorithms can cater for noncircularity of the input a necessary feature in applications for modelling real-world data. It is furthermore shown how the framework can be constrained to incorporate nonnegativity, helping generate results which allow a more realistic interpretation. Simulations illustrate the usefulness and enhanced accuracy for modelling synthetic data and a mixture of acoustic stimuli.
David Looney, Danilo P. Mandic
ICASSP2
2011 A collaborative filtering approach for quasi-brain-death EEG analysis
abstract
A novel method to evaluate the statistical significance differences between the groups of coma and brain death patients is presented. This is achieved based on the electroencephalogram (EEG) and by using a collaborative filtering structure with the least mean square (LMS) and least mean phase (LMP) adaptive filters. By virtue of a complex-valued representation of pair-wise EEG signals, the evolution of the mixing parameter is used as an indicator of the fundamental amplitude-phase relationships of EEG recordings. Simulations illustrate the suitability of this approach to differentiate between the coma and quasi-brain-death states.
Yili Xia, Ling Li 0010, Jianting Cao, Martin Golz, Danilo P. Mandic
ICASSP5
2011 A class of fast quaternion valued variable stepsize stochastic gradient learning algorithms for vector sensor processes
abstract
We introduce a class of gradient adaptive stepsize algorithms for quaternion valued adaptive filtering based on three- and four-dimensional vector sensors. This equips the recently introduced quaternion least mean square (QLMS) algorithm with enhanced tracking ability and enables it to be more responsive to dynamically changing environments, while maintaining its desired characteristics of catering for large dynamical differences and coupling between signal components. For generality, the analysis is performed for the widely linear signal model, which by virtue of accounting for signal noncircularity, is optimal in the mean squared error (MSE) sense for both second order circular (proper) and noncircular (improper) processes. The widely linear QLMS (WL-QLMS) employing the proposed adaptive stepsize modifications is shown to provide enhanced performance for both synthetic and real world quaternion valued signals. Simulations include signals with drastically different component dynamics, such as four dimensional quaternion comprising three dimensional turbulent wind and air temperature for renewable energy applications.
Mingxuan Wang, Clive Cheong Took, Danilo P. Mandic
IJCNN3
2011 Widely linear adaptive frequency estimation in three-phase power systems under unbalanced voltage sag conditions
abstract
A new framework for the estimation of the instantaneous frequency in a three-phase power system is proposed. It is first illustrated that the complex-valued signal, obtained by the αβ transformation of three-phase power signals under unbalanced voltage sag conditions, is second order noncircular, for which standard complex adaptive estimators are suboptimal. To cater for second order noncircularity, an adaptive widely linear estimator based on the augmented complex least mean square (ACLMS) algorithm is proposed, and the analysis shows that this allows for optimal linear adaptive estimation for the generality of system conditions (both balanced and unbalanced). The enhanced robustness over the standard CLMS is illustrated by simulations on both synthetic and real-world voltage sags.
Yili Xia, Scott C. Douglas, Danilo P. Mandic
IJCNN3
2011 Multilinear Subspace Regression: An Orthogonal Tensor Decomposition Approach
abstract
A multilinear subspace regression model based on so called latent variable decomposition is introduced. Unlike standard regression methods which typically employ matrix (2D) data representations followed by vector subspace transformations, the proposed approach uses tensor subspace transformations to model common latent variables across both the independent and dependent data. The proposed approach aims to maximize the correlation between the so derived latent variables and is shown to be suitable for the prediction of multidimensional dependent data from multidimensional independent data, where for the estimation of the latent variables we introduce an algorithm based on Multilinear Singular Value Decomposition (MSVD) on a specially defined cross-covariance tensor. It is next shown that in this way we are also able to unify the existing Partial Least Squares (PLS) and N-way PLS regression algorithms within the same framework. Simulations on benchmark synthetic data confirm the advantages of the proposed approach, in terms of its predictive ability and robustness, especially for small sample sizes. The potential of the proposed technique is further illustrated on a real world task of the decoding of human intracranial electrocorticogram (ECoG) from a simultaneously recorded scalp electroencephalograph (EEG).
Qibin Zhao, Cesar F. Caiafa, Danilo P. Mandic, Liqing Zhang 0001, Tonio Ball, Andreas Schulze-Bonhage, Andrzej Cichocki
NIPS3
2011 The complex local mean decomposition
David Looney, Marc M. Van Hulle, Danilo P. Mandic
Neurocomputing4
2011 Augmented second-order statistics of quaternion random signals
Clive Cheong Took, Danilo P. Mandic
Signal Process.2
2011 A Widely Linear Complex Unscented Kalman Filter
abstract
Conventional complex valued signal processing algorithms assume rotation invariant (circular) signal distributions, and are thus suboptimal for real world processes which exhibit rotation dependent distributions (noncircular). In nonlinear sequential state space estimation, noncircularity can arise from the data, state transition model, and state and observation noises. We provide further insight by revisiting the augmented complex unscented Kalman filter (ACUKF) and illuminating its operation in such scenarios. The analysis establishes a relationship between the estimation error and the degree of second order noncircularity (improperness) in the system for the conventional complex unscented Kalman filter (CUKF), and is supported by simulations on both synthetic and real world proper and improper signals.
Dahir H. Dini, Danilo P. Mandic, Simon J. Julier
IEEE Signal Process. Lett.2
2011 A Quaternion Gradient Operator and Its Applications
abstract
Real functions of quaternion variables are typical cost functions in quaternion valued statistical signal processing, however, standard differentiability conditions in the quaternion domain do not permit direct calculation of their gradients. To this end, based on the isomorphism with real vectors and the use of quaternion involutions, we introduce the HR calculus as a convenient way to calculate derivatives of such functions. It is shown that the maximum change of the gradient is in the direction of the conjugate gradient, which conforms with the corresponding solution in the complex domain. Examples in some typical gradient based optimization settings support the result.
Danilo P. Mandic, Cyrus Jahanchahi, Clive Cheong Took
IEEE Signal Process. Lett.1
2011 An Adaptive Diffusion Augmented CLMS Algorithm for Distributed Filtering of Noncircular Complex Signals
abstract
An adaptive diffusion augmented complex least mean square (D-ACLMS) algorithm for collaborative processing of the generality of complex signals over distributed networks is proposed. The algorithm enables the estimation of both second order circular (proper) and noncircular (improper) signals within a unified framework of augmented complex statistics. The analysis shows that the performance advantage of the widely linear D-ACLMS over the strictly linear D-CLMS increases with the degree of noncircularity while maintaining similar performance for proper data. Simulations on both synthetic benchmark and real world noncircular data support the approach.
Yili Xia, Danilo P. Mandic, Ali H. Sayed
IEEE Signal Process. Lett.2
2011 Fast Independent Component Analysis Algorithm for Quaternion Valued Signals
abstract
An extension of the fast independent component analysis algorithm is proposed for the blind separation of both Q-proper and Q-improper quaternion-valued signals. This is achieved by maximizing a negentropy-based cost function, and is derived rigorously using the recently developed HR calculus in order to implement Newton optimization in the augmented quaternion statistics framework. It is shown that the use of augmented statistics and the associated widely linear modeling provides theoretical and practical advantages when dealing with general quaternion signals with noncircular (rotation-dependent) distributions. Simulations using both benchmark and real-world quaternion-valued signals support the approach.
Soroush Javidi, Clive Cheong Took, Danilo P. Mandic
IEEE Trans. Neural Networks3
2011 Quaternion-Valued Nonlinear Adaptive Filtering
abstract
A class of nonlinear quaternion-valued adaptive filtering algorithms is proposed based on locally analytic nonlinear activation functions. To circumvent the stringent standard analyticity conditions which are prohibitive to the development of nonlinear adaptive quaternion-valued estimation models, we use the fact that stochastic gradient learning algorithms require only local analyticity at the operating point in the estimation space. It is shown that the quaternion-valued exponential function is locally analytic, and, since local analyticity extends to polynomials, products, and ratios, we show that a class of transcendental nonlinear functions can serve as activation functions in nonlinear and neural adaptive models. This provides a unifying framework for the derivation of gradient-based learning algorithms in the quaternion domain, and the derived algorithms are shown to have the same generic form as their real- and complex-valued counterparts. To make such models second-order optimal for the generality of quaternion signals (both circular and noncircular), we use recent developments in augmented quaternion statistics to introduce widely linear versions of the proposed nonlinear adaptive quaternion valued filters. This allows full exploitation of second-order information in the data, contained both in the covariance and pseudocovariances to cater rigorously for second-order noncircularity (improperness), and the corresponding power mismatch in the signal components. Simulations over a range of circular and noncircular synthetic processes and a real world 3-D noncircular wind signal support the approach.
Bukhari Che Ujang, Clive Cheong Took, Danilo P. Mandic
IEEE Trans. Neural Networks3
2011 An Augmented Echo State Network for Nonlinear Adaptive Filtering of Complex Noncircular Signals
abstract
A novel complex echo state network (ESN), utilizing full second-order statistical information in the complex domain, is introduced. This is achieved through the use of the so-called augmented complex statistics, thus making complex ESNs suitable for processing the generality of complex-valued signals, both second-order circular (proper) and noncircular (improper). Next, in order to deal with nonstationary processes with large nonlinear dynamics, a nonlinear readout layer is introduced and is further equipped with an adaptive amplitude of the nonlinearity. This combination of augmented complex statistics and enhanced adaptivity within ESNs also facilitates the processing of bivariate signals with strong component correlations. Simulations in the prediction setting on both circular and noncircular synthetic benchmark processes and real-world noncircular and nonstationary wind signals support the analysis.
Yili Xia, Beth Jelfs, Marc M. Van Hulle, José C. Príncipe, Danilo P. Mandic
IEEE Trans. Neural Networks5
2010 Image fusion based on Fast and Adaptive Bidimensional Empirical Mode Decomposition
Mosabber Uddin Ahmed, Danilo P. Mandic
FUSION2
2010 Performance analysis of the conventional complex LMS and augmented complex LMS algorithms
abstract
Recently, the augmented complex LMS (ACLMS) algorithm has been proposed for modeling complex-valued signal relationships in which a widely-linear model can be more appropriate. It is not clear, however, how the behavior of ACLMS differs from that of the conventional complex LMS (CCLMS) algorithm. In this paper, we leverage a recently-developed analysis for the complex LMS algorithm to illuminate the performance relationships between the ACLMS and CCLMS algorithms. Our analysis shows that the ACLMS algorithm can potentially achieve a lower steady-state mean-squared error as compared to that of CCLMS, but the convergence speed of ACLMS is slowed in the presence of highly non-circular complex-valued input signals. An adaptive beamforming example indicates the utility of the results.
Scott C. Douglas, Danilo P. Mandic
ICASSP2
2010 An Auditory Oddball Based Brain-Computer Interface System Using Multivariate EMD
Qi-Wei Shi, Jianting Cao, Danilo P. Mandic, Toshihisa Tanaka 0001, Tomasz M. Rutkowski, Rubin Wang
ICIC (2)4
2010 On HR calculus, quaternion valued stochastic gradient, and adaptive three dimensional wind forecasting
abstract
Short term forecasting of wind field in the quaternion domain is addressed. This is achieved by casting the three components of wind speed (two horizontal and a vertical) into a pure quaternion and adding air temperature as a scalar component, to form the full quaternion. First, HR calculus is introduced in order to provide a unifying framework for the calculation of the derivatives of both analytic quaternion valued functions and real functions of quaternion variables, such as the standard cost function (error power). The analysis shows that the maximum change in the gradient is in the direction of the conjugate of the weight vector, conforming with the gradient calculation in the complex domain. For rigour, we also illustrate that the widely linear model is required in order to capture full second order information within three- and four-dimensional quaternion valued signals. The so established framework is used to illustrate a convenient way to derive the recently introduced quaternion least mean square (QLMS) and the widely linear QLMS (WL-QLMS). Simulations on short term prediction of real world wind signals support the approach.
Cyrus Jahanchahi, Clive Cheong Took, Danilo P. Mandic
IJCNN3
2010 Towards estimating selective auditory attention from EEG using a novel time-frequency-synchronisation framework
abstract
An original experimental design is combined with a novel signal processing approach so as to provide cognitive clues in the study of auditory scene analysis and in the design of auditory brain computer interfaces. Volunteers attended a single auditory stimulus in a perceptually complex auditory environment of speech and music, wherein the experiment aim was to estimate the attended stimulus from recorded electroencephalogram (EEG). Unlike previous studies, the complex nature of the auditory environment does not allow for straightforward analysis that exploits convenient properties of the stimuli. To provide insight, synchronised neuronal activity was analysed within a novel signal processing framework that models energy and phase dynamics independently using empirical mode decomposition. By design, the proposed approach caters for higher order information and is suitable for nonstationary data, both critical properties in the analysis of cognitive activity. The proposed methodology achieved a median classification accuracy of 71% in a series of selective attention experiments with several volunteers.
David Looney, Yili Xia, Preben Kidmose, Michael Ungstrup, Danilo P. Mandic
IJCNN6
2010 Quadrivariate Empirical Mode Decomposition
abstract
We introduce a quadrivariate extension of Empirical Mode Decomposition (EMD) algorithm, termed QEMD, as a tool for the time-frequency analysis of nonlinear and non-stationary signals consisting of up to four channels. The local mean estimation of the quadrivariate signal is based on taking real-valued projections of the input in different directions in a multidimensional space where the signal resides. To this end, the set of direction vectors is generated on 3-sphere (residing in 4D space) via the low-discrepancy Hammersley sequence. It has also been shown that the resulting set of vectors is more uniformly distributed on a 3-sphere as compared to that generated by a uniform angular coordinate system. The ability of QEMD to extract common modes within multichannel data is demonstrated by simulations on both synthetic and real-world signals.
Naveed ur Rehman, Danilo P. Mandic
IJCNN2
2010 Quaternion-valued short term forecasting of wind profile
abstract
This work presents novel methodology for the simultaneous modelling and forecasting of three-dimensional (3D) wind fields. This is achieved based on a quaternion domain wind model, which naturally accounts for the coupling between the dimensions of the 3D wind field. The proposed quaternion valued processing also facilitates the fusion of external atmospheric parameters, such as air temperature, exhibiting more degrees of freedom and enhanced accuracy. The quaternion least mean square (QLMS) algorithm and its variants are used for short term adaptive forecasting, and a rigorous comparative study with the corresponding algorithms in ℝ4is performed. Simulations for different wind regimes and over a range of prediction horizons support the approach.
Clive Cheong Took, Danilo P. Mandic, Kazuyuki Aihara
IJCNN2
2010 Identification of improper processes by variable tap-length complex-valued adaptive filters
abstract
Variable tap-length is introduced into complex-valued adaptive filters in order to provide an additional degree of freedom, enhance tracking ability, and provide data-adaptive optimal modelling. This is achieved by extending the fractional tap-length (FT) algorithm from the real domain ℝ and by accounting for some special properties of the complex domain ℂ. For generality, the augmented least mean square (ACLMS) and augmented complex nonlinear gradient descent (ACNGD) are equipped with the variable tap-length in order to cater for both the second order circular and noncircular signals. Simulations on model order selection and the identification of the noncircular nature of complex data support the approach.
Bukhari Che Ujang, Clive Cheong Took, Danilo P. Mandic
IJCNN3
2010 Split quaternion nonlinear adaptive filtering
Bukhari Che Ujang, Clive Cheong Took, Danilo P. Mandic
Neural Networks3
2010 An augmented affine projection algorithm for the filtering of noncircular complex signals
Yili Xia, Clive Cheong Took, Danilo P. Mandic
Signal Process.3
2009 Applications of complex augmented kernels to wind profile prediction
abstract
This paper combines complex signal processing with kernel methods for applications in wind prediction. Specifically, we consider developing least squares kernel algorithms for both complex data and augmented complex data. The augmented complex kernel algorithms have advantages over complex kernel algorithms in both the areas of performance and complexity. Use of kernels also allow implementation of nonlinear algorithms by working in the dual space. We apply our algorithm to wind series time prediction and show that our augmented complex algorithms outperform other complex least square algorithms.
Anthony Kuh, Danilo P. Mandic
ICASSP2
2009 Duality between widely linear and dual channel adaptive filtering
abstract
We address the duality between adaptive filtering in C and R2and provide a comparison between the well understood dual channel real valued least mean square (DCRLMS) algorithm in R2and the corresponding algorithms in C. These include the complex LMS (CLMS) and the recently introduced augmented CLMS (ACLMS), a widely linear algorithm designed for the processing of noncircular complex valued signals. The analysis shows that the standard CLMS and DCRLMS in general provide different adaptive filtering solutions, whereas the ACLMS and DCRLMS are isomoprhic and can be made equivalent. The analysis is supported by simulations on noncircular real world signals.
Danilo P. Mandic, Susanne Still, Scott C. Douglas
ICASSP1
2009 Qualitative analysis of rotational modes within three dimensional empirical mode decomposition
abstract
An analysis of quaternion-valued intrinsic mode functions (IMFs) within three dimensional empirical mode decomposition is presented. This is achieved by using the delay vector variance (DVV) method, which examines the signal predictability in phase space to assess the determinism and nonlinearity within the signal. The study illustrates that the contribution of the first few IMFs contain information related to the stochastic/nonlinear signal nature, whereas the lower order IMFs are largely deterministic. The analysis is supported by simulation results on a quaternion signal composed of linear/nonlinear benchmark signals and on real world wind data.
Naveed ur Rehman, Danilo P. Mandic
ICASSP2
2009 Multichannel spectral pattern separation - An EEG processing application -
abstract
A problem of information separation in multichannel recordings is important in engineering applications such as brain computer/machine interfaces (BCI/BMI). Whereas this problem is not entirely new, engineering approaches connecting the mental states of humans and the observed electroencephalography (EEG) recordings are still in their infancy, mostly due to problems with electrophysiological denoising. The electrophysiological signals captured in form of the EEG carry brain activity in form of the neurophysiological components which are usually embedded in much higher power electrical muscle activity components (electromyography - EMG; electrooculography - EOG; etc.). In this paper we present an approach to remove muscular interference caused by eye-movements from EEG recorded during auditory experiments in an eight channel recording setting. This is achieved by analyzing the correlation of the oscillatory modes within a multichannel signal in the Hilbert domain. Simulations in a real world auditory BCI setting support the analysis.
Tomasz M. Rutkowski, Andrzej Cichocki, Toshihisa Tanaka 0001, Danilo P. Mandic, Jianting Cao, Anca L. Ralescu
ICASSP4
2009 Study of the quaternion LMS and four-channel LMS algorithms
abstract
The recently proposed quaternion least-mean-square (QLMS) algorithm for adaptive filtering of three- and four-dimensional signals has been analysed in the context of multi-step ahead prediction. For rigour, the relationship between multichannel LMS (MLMS) and QLMS is examined, and their differences are highlighted. This is achieved both in terms of the input-output relationship and in terms of the dynamics of weight updates. The convergence of QLMS is investigated and stability bounds confirm that QLMS and MLMS are fundamentally different. Simulations on both synthetic and real world multidimensional signals support the analysis.
Clive Cheong Took, Danilo P. Mandic, Jacob Benesty
ICASSP2
2009 A split quaternion nonlinear adaptive filter
abstract
A split quaternion learning algorithm for the training of nonlinear finite impulse response filters for the modelling of hypercomplex signals is proposed. A rigorous derivation takes into account the non-commutativity of the quaternion product, an aspect not taken into account in the existing nonlinear architectures, such as the quaternion multilayer perceptron (QMLP). It is shown that the additional information present within the proposed algorithm provides an improved performance over QMLP. Simulation on both benchmark and real-world signals support the approach.
Bukhari Che Ujang, Clive Cheong Took, Aleksandar Kavcic, Danilo P. Mandic
ICASSP4
2009 Complex valued recurrent neural networks for noncircular complex signals
abstract
This paper uses new developments in the statistics of complex variable and recent results on the duality between the bivariate and complex calculus to provide a unified design of complex valued temporal neural networks. For generality, the case of recurrent neural networks is addressed in detail, as they simplify into feedforward networks upon cancellation of the feedback. The use of CopfRopf calculus provides a convenient framework for the calculation of gradients of real functions of complex variables (cost functions) which do not obey the Cauchy-Riemann conditions. Further, the analysis is based on so called augmented complex statistics, to provide a rigorous treatment of complex noncircularity and nonlinearity, thus avoiding the deficiencies inherent in several mathematical shortcuts typically used in the treatment of complex random signals. The complex models addressed in this work, are based on widely linear nonlinear autoregressive moving average (NARMA) models and are shown to be suitable for processing the generality of complex signals, both second order circular (proper) and noncircular (improper).
Danilo P. Mandic
IJCNN1
2009 Upper bounds on the capacities of non-controllable finite-state channels using dynamic programming methods
abstract
A non-controllable finite-state channel (FSC) is a finite-state channel in which the user can't control channel states. That is, the channel state of a non-controllable FSC evolves freely according to an uncontrollable probability law. Thus far, good upper bounds on capacities of general non-controllable FSCs remain unknown. Here we develop upper bounds that use delayed feedback and delayed state information, and propose dynamic programming methods to numerically evaluate the bounds.
Xiujie Huang, Aleksandar Kavcic, Xiao Ma 0001, Danilo P. Mandic
ISIT4
2008 Qualitative assessment of intrinsic mode functions of empirical mode decomposition
abstract
The 'empirical mode decomposition' (EMD) method has been recently proposed to deal with nonlinear and non- stationary data, which decomposes signals into 'well-behaved' intrinsic mode functions (IMFs). An assessment on the qualitative performance of the EMD method in terms of the degree of signal nature preservation of individual IMF is provided. This is archived by means of the recently proposed signal characterisation method, based upon examining the signal predicability in phase space. It is shown that the first IMF always performs best in terms of signal nature preserving. Simulation results on both linear and nonlinear benchmark signals support the analysis.
Mo Chen 0006, Danilo P. Mandic, Preben Kidmose, Michael Ungstrup
ICASSP2
2008 Cascaded approach for microsleep data extraction
abstract
The noisy component extraction (NoiCE) algorithm is proposed to blind-extract noisy signals. This is achieved based on a combination of blind extraction structure and a cascaded nonlinear adaptive estimation. Although we use the concept of sequential blind extraction of sources and independent component analysis (ICA), we do not assume that sources are statistically independent. In fact, we show that the proposed cascaded nonlinear filter can be used to extract a signal (a single signal each time) from their noisy mixtures. Computer simulations confirm the validity and performance of the proposed algorithm in noisy microsleep events.
Wai Yie Leong, Danilo P. Mandic
ICASSP2
2008 A machine learning enhanced empirical mode decomposition
abstract
Empirical mode decomposition (EMD) is a fully data driven method for decomposing signals into a set of AM-FM components known as intrinsic mode functions (IMFs). Despite its usefulness in the analysis of real world signals, the process is rather deterministic and sensitive to parameters such as local envelope estimation. A combination of EMD and machine learning is proposed which provides an algorithm that is more robust to EMD parameters. In addition, the proposed extension is fully adaptive and facilitates the "data fusion via fission" mode of operation. The derivation and analysis of the proposed framework is supported with simulations in denoising and prediction applications.
David Looney, Danilo P. Mandic
ICASSP2
2008 Online tracking of the degree of nonlinearity within complex signals
abstract
A novel method for online tracking of the changes in the non- linearity within complex-valued signals is introduced. This is achieved by a collaborative adaptive signal processing approach by means of a hybrid filter. By tracking the dynamics of the adaptive mixing parameter within the employed hybrid filtering architecture, we show that it is possible to quantify the degree of nonlinearity within complex-valued data. Simulations on both benchmark and real world data support the approach.
Danilo P. Mandic, Phebe Vayanos, Soroush Javidi, Beth Jelfs, Kazuyuki Aihara
ICASSP1
2008 EMD Approach to Multichannel EEG Data - The Amplitude and Phase Synchrony Analysis Technique
Tomasz M. Rutkowski, Danilo P. Mandic, Andrzej Cichocki, Andrzej W. Przybyszewski
ICIC (1)2
2008 Clustering of Spectral Patterns Based on EMD Components of EEG Channels with Applications to Neurophysiological Signals Separation
Tomasz M. Rutkowski, Andrzej Cichocki, Toshihisa Tanaka 0001, Anca L. Ralescu, Danilo P. Mandic
ICONIP (1)5
2008 Online Detection of the Modality of Complex-Valued Real World Signals
abstract
A novel method for the online detection of the modality of complex-valued nonlinear and nonstationary signals is introduced. This is achieved using a convex combination of complex nonlinear adaptive filters with different transient characteristics. To facilitate the online mode of operation, the convex mixing parameter lambda within the proposed architecture is made gradient adaptive. Our focus is on the most important aspect of complex nonlinear modeling, that is, the identification of the split-complex and fully-complex nature of the signal in hand. The algorithms derived are robust and capable of tracking the changes in the modality of both benchmark and real world radar and wind complex vector fields.
Danilo P. Mandic, Phebe Vayanos, Mo Chen 0006, Su Lee Goh
Int. J. Neural Syst.1
2008 Advances in blind signal processing
Deniz Erdogmus, Danilo P. Mandic, Toshihisa Tanaka 0001
Neurocomputing2
2008 Blind source extraction: Standard approaches and extensions to noisy and post-nonlinear mixing
Wai Yie Leong, Wei Liu 0001, Danilo P. Mandic
Neurocomputing3
2008 A Homomorphic Neural Network for Modeling and Prediction
abstract
A homomorphic feedforward network (HFFN) for nonlinear adaptive filtering is introduced. This is achieved by a two-layer feedforward architecture with an exponential hidden layer and logarithmic preprocessing step. This way, the overall input-output relationship can be seen as a generalized Volterra model, or as a bank of homomorphic filters. Gradient-based learning for this architecture is introduced, together with some practical issues related to the choice of optimal learning parameters and weight initialization. The performance and convergence speed are verified by analysis and extensive simulations. For rigor, the simulations are conducted on artificial and real-life data, and the performances are compared against those obtained by a sigmoidal feedforward network (FFN) with identical topology. The proposed HFFN proved to be a viable alternative to FFNs, especially in the critical case of online learning on small- and medium-scale data sets.
Maciej Pedzisz, Danilo P. Mandic
Neural Comput.2
2008 Preface
Danilo P. Mandic, Wlodzislaw Duch
Neural Networks1
2008 An Assessment of Qualitative Performance of Machine Learning Architectures: Modular Feedback Networks
abstract
A framework for the assessment of qualitative performance of machine learning architectures is proposed. For generality, the analysis is provided for the modular nonlinear pipelined recurrent neural network (PRNN) architecture. This is supported by a sensitivity analysis, which is achieved based upon the prediction performance with respect to changes in the nature of the processed signal and by utilizing the recently introduced delay vector variance (DVV) method for phase space signal characterization. Comprehensive simulations combining the quantitative and qualitative analysis on both linear and nonlinear signals suggest that better quantitative prediction performance may need to be traded in order to preserve the nature of the processed signal, especially where the signal nature is of primary importance (biomedical applications).
Mo Chen 0006, Temujin Gautama, Danilo P. Mandic
IEEE Trans. Neural Networks3
2007 Rotation Invariant Complex Empirical Mode Decomposition
abstract
A new method to extend the empirical mode decomposition (EMD) into the complex domain is proposed. Unlike the existing method for EMD in the complex domain, this is achieved in a generic way so that the mathematical development of this method mirrors the algorithm defined for EMD in the real domain. The so derived intrinsic mode functions (IMFs) are complex by design and are shown to provide a consistent framework for handling both real and complex data. The simulations on real world complex-valued signals illustrate the applications of the technique.
Temujin Gautama, Toshihisa Tanaka 0001, Danilo P. Mandic
ICASSP (3)4
2007 Blind Extraction of Noisy Events using Nonlinear Predictor
abstract
Existing blind source extraction (BSE) methods are limited to noise-free mixtures, which is not realistic. We therefore address this issue and propose an algorithm based on the normalised kurtosis and a nonlinear predictor within the BSE structure, which makes this class of algorithms suitable for noisy environments, a typical situation in practice. Based on a rigorous analysis of the existing BSE methods we also propose a new optimisation paradigm which aims at minimising the normalised mean square prediction error (MSPE). This makes redundant the need for preprocessing or orthogonality transform. Simulation results are provided which confirm the validity of the theoretical results and demonstrate the performance of the derived algorithms in noisy mixing environments.
Wai Yie Leong, Danilo P. Mandic, Wei Liu 0001
ICASSP (2)2
2007 Collaborative Adaptive Learning using Hybrid Filters
abstract
A novel stable and robust algorithm for training of finite impulse response adaptive filters is proposed. This is achieved based on a convex combination of the least mean square (LMS) and a recently proposed generalised normalised gradient descent (GNGD) algorithm. In this way, the desirable fast convergence and stability of GNGD is combined with the robustness and small steady state misadjustment of LMS. Simulations on linear and nonlinear signals in the prediction setting support the analysis.
Danilo P. Mandic, Phebe Vayanos, Christos Boukis, Beth Jelfs, Su Lee Goh, Temujin Gautama, Tomasz M. Rutkowski
ICASSP (3)1
2007 Noisy Component Extraction (Noice)
abstract
Existing blind source extraction (BSE) methods are limited to noise-free mixtures, which is not realistic. Based on a rigorous analysis of the existing BSE methods, we address the problem of noisy component extraction (NoiCE) which provides BSE in the presence of noise. As a byproduct in BSE after deflation, we may also obtain the asymptotic identification of the a priori unknown observation noise disturbance. By yielding an asymptotically efficient filter in the presence of an unknown observation noise, our approach may also be viewed as a robust approach to noisy component extraction. Simulation results are provided which confirm the validity of the theoretical results and demonstrate the performance of the derived algorithms in noisy mixing environments.
Wai Yie Leong, Danilo P. Mandic
ISCAS2
2007 Algorithms for BER-Constrained Variable-Length Equalizers driven by Channel Response Knowledge over Frequency-Selective Radio Channel
abstract
In mobile radio systems, transmission conditions actually encounter large variations depending on the effective environmental configurations. Training sequences are periodically inserted into the transmitted messages so that the receiver can estimate the channel response. Based on this knowledge, we study the problem of adapting the equalizer filter length to a given channel impulse response under bit error rate constraint. In this paper, both linear equalizer (LE) and decision feedback equalizer (DFE) are studied. Accurate control of the equalizer length can improve the system's performance while reducing the handset power consumption or reducing software load in a software radio context. Besides, in a multi-service system, it allows to manage variable quality of service requirements. Optimal and suboptimal criteria for determining the equalizer length are discussed in order to optimize the overall complexity of the receiver including training phase and decoding phases.
Armelle Wautier, Lionel Husson, Ionut-Dan Plai, Danilo P. Mandic
VTC Spring4
2007 An Augmented Extended Kalman Filter Algorithm for Complex-Valued Recurrent Neural Networks
abstract
An augmented complex-valued extended Kalman filter (ACEKF) algorithm for the class of nonlinear adaptive filters realized as fully connected recurrent neural networks is introduced. This is achieved based on some recent developments in the so-called augmented complex statistics and the use of general fully complex nonlinear activation functions within the neurons. This makes the ACEKF suitable for processing general complex-valued nonlinear and nonstationary signals and also bivariate signals with strong component correlations. Simulations on benchmark and real-world complex-valued signals support the approach.
Su Lee Goh, Danilo P. Mandic
Neural Comput.2
2007 An augmented CRTRL for complex-valued recurrent neural networks
Su Lee Goh, Danilo P. Mandic
Neural Networks2
2007 Biometrics from Brain Electrical Activity: A Machine Learning Approach
abstract
The potential of brain electrical activity generated as a response to a visual stimulus is examined in the context of the identification of individuals. Specifically, a framework for the Visual Evoked Potential (VEP)-based biometrics is established, whereby energy features of the gamma band within VEP signals were of particular interest. A rigorous analysis is conducted which unifies and extends results from our previous studies, in particular, with respect to 1) increased bandwidth, 2) spatial averaging, 3) more robust power spectrum features, and 4) improved classification accuracy. Simulation results on a large group of subject support the analysis.
Ramaswamy Palaniappan, Danilo P. Mandic
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Complex Empirical Mode Decomposition
abstract
A method for the empirical mode decomposition (EMD) of complex-valued data is proposed. This is achieved based on the filter bank interpretation of the EMD mapping and by making use of the relationship between the positive and negative frequency component of the Fourier spectrum. The so-generated intrinsic mode functions (IMFs) are complex-valued, which facilitates the extension of the standard EMD to the complex domain. The analysis is supported by simulations on both synthetic and real-world complex-valued signals
Toshihisa Tanaka 0001, Danilo P. Mandic
IEEE Signal Process. Lett.2
2007 Stochastic Gradient-Adaptive Complex-Valued Nonlinear Neural Adaptive Filters With a Gradient-Adaptive Step Size
abstract
A class of variable step-size learning algorithms for complex-valued nonlinear adaptive finite impulse response (FIR) filters is proposed. To achieve this, first a general complex-valued nonlinear gradient-descent (CNGD) algorithm with a fully complex nonlinear activation function is derived. To improve the convergence and robustness of CNGD, we further introduce a gradient-adaptive step size to give a class of variable step-size CNGD (VSCNGD) algorithms. The analysis and simulations show the proposed class of algorithms exhibiting fast convergence and being able to track nonlinear and nonstationary complex-valued signals. To support the derivation, an analysis of stability and computational complexity of the proposed algorithms is provided. Simulations on colored, nonlinear, and real-world complex-valued signals support the analysis.
Su Lee Goh, Danilo P. Mandic
IEEE Trans. Neural Networks2
2007 Analysis and Online Realization of the CCA Approach for Blind Source Separation
abstract
A critical analysis of the canonical correlation analysis (CCA) approach in blind source separation (BSS) is provided. It is proved that by maximizing the autocorrelation functions of the recovered signals we can separate the source signals successfully. It is further shown that the CCA approach represents the same class of generalized eigenvalue decomposition (GEVD) problems as the matrix pencil method. Finally, online realizations of the CCA approach are discussed with a linear-predictor-based algorithm studied as an example.
Wei Liu 0001, Danilo P. Mandic, Andrzej Cichocki
IEEE Trans. Neural Networks2
2006 An Augmented Extended Kalman Filter Algorithm for Complex-Valued Recurrent Neural Networks
abstract
An augmented complex-valued Extended Kalman Filter (ACEKF) algorithm for the class of nonlinear adaptive filters realised as fully connected recurrent neural networks (FCRNNs) is introduced. The algorithm is derived based on the recent developments in augmented complex statistics, and the Jacobian matrix within the ACEKF algorithm is computed using a general fully complex real time recurrent learning (CRTRL) algorithm. This makes ACEKF suitable for processing general complex-valued nonlinear and nonstationary signals and bivariate signals with strong component correlations. Simulations on benchmark and real-world complexvalued signals support the approach.
Su Lee Goh, Danilo P. Mandic
ICASSP (5)2
2006 Sequential Detection Using Least Squares Temporal Difference Methods
abstract
This paper considers sequential detection problems where we learn from sets of training sequences. The sufficient statistics can be learned quickly using a least squares temporal difference (TD) learning algorithm. This algorithm converges much quicker than previously applied TD learning algorithms. The algorithm can easily be implemented in an on-line manner and can also be applied to more complicated decentralized detection problems
Anthony Kuh, Danilo P. Mandic
ICASSP (5)2
2006 A Normalised Kurtosis Based Blind Source Extraction Algorithm for Noisy Mixtures
abstract
We introduce an algorithm for blind source extraction (BSE) of independent sources in the presence of noise, without the need for initial prewhitening, for which the normalised kurtosis is used within the cost function. Unlike the previously proposed methods designed for noise-free mixtures, which is not realistic in practical applications, we address BSE for noisy mixtures and propose a novel cost function which caters for the effects of noise. The proposed method is justified by rigorous analysis and supported by simulations
Wei Liu 0001, Danilo P. Mandic
ICASSP (5)2
2006 An analysis of the CCA approach for blind source separation and its adaptive realization
abstract
An analysis of the canonical correlation analysis (CCA) approach in blind source separation is provided. In particular, it is proved that by maximizing the autocorrelation functions of the recovered signals we can separate the source signals successfully. We show that the CCA approach represents the same generalised eigenvalue decomposition problem introduced in the matrix pencil method. Finally, an adaptive blind source extraction (BSE) algorithm is derived as an online realisation of the CCA approach. Simulation results verify the proposed approach
Wei Liu 0001, Danilo P. Mandic, Andrzej Cichocki
ISCAS2
2006 Blind source extraction of instantaneous noisy mixtures using a linear predictor
abstract
The blind source extraction (BSE) problem for noisy measurements is addressed using the linear predictor method. Based on a previously proposed method for the noise-free case, we propose a cost function with the effect of noise removed. Two adaptive algorithms are next introduced, one of which is based on minimisation of the normalised mean square prediction error (MSPE), whereas the other minimises the MSPE using prewhitening followed by regularisation of the demixing vector. The successful operation of these algorithms requires the knowledge of the correlation matrix of noise
Wei Liu 0001, Danilo P. Mandic, Andrzej Cichocki
ISCAS2
2006 An Online Method for Detecting Nonlinearity Within a Signal
Beth Jelfs, Phebe Vayanos, Mo Chen 0006, Su Lee Goh, Christos Boukis, Temujin Gautama, Tomasz M. Rutkowski, Tony Kuh, Danilo P. Mandic
KES (3)9
2006 Sensor Network Localization Using Least Squares Kernel Regression
Anthony Kuh, Chaopin Zhu, Danilo P. Mandic
KES (3)3
2006 Auditory Feedback for Brain Computer Interface Management - An EEG Data Sonification Approach
Tomasz M. Rutkowski, François B. Vialatte, Andrzej Cichocki, Danilo P. Mandic, Allan Kardec Barros
KES (3)4
2006 A Flexible Method for Envelope Estimation in Empirical Mode Decomposition
Yoshikazu Washizawa, Toshihisa Tanaka 0001, Danilo P. Mandic, Andrzej Cichocki
KES (3)3
2006 A normalised kurtosis-based algorithm for blind source extraction from noisy measurements
Wei Liu 0001, Danilo P. Mandic
Signal Process.2
2006 A novel algorithm for the adaptation of the pole of Laguerre filters
abstract
This letter proposes a novel stochastic gradient algorithm for the online adaptation of the pole position in Laguerre filters. The proposed algorithm exploits the inherent relationship between the values of the filter coefficients and the value of the Laguerre pole. This leads to an unbiased solution and, hence, a more accurate estimate of the error gradient. Simulations in a system identification setting support the analysis
Christos Boukis, Danilo P. Mandic, Anthony G. Constantinides, Lazaros Polymenakos
IEEE Signal Process. Lett.2
2005 Data Fusion for Modern Engineering Applications: An Overview
Danilo P. Mandic, Dragan Obradovic, Anthony Kuh, Tülay Adali, Udo Trutschel, Martin Golz, Philippe De Wilde, Javier A. Barria, Anthony G. Constantinides, Jonathon A. Chambers
ICANN (2)1
2005 Energy of Brain Potentials Evoked During Visual Stimulus: A New Biometric?
Ramaswamy Palaniappan, Danilo P. Mandic
ICANN (2)2
2005 Communicative Interactivity - A Multimodal Communicative Situation Classification Approach
Tomasz M. Rutkowski, Danilo P. Mandic
ICANN (2)2
2005 Fusion of State Space and Frequency- Domain Features for Improved Microsleep Detection
David Sommer, Mo Chen 0006, Martin Golz, Udo Trutschel, Danilo P. Mandic
ICANN (2)5
2005 On nonlinear modular neural filters
abstract
An assessment of the performance of the pipelined recurrent neural network (PRNN) is provided from two aspects, a quantitative one based on the prediction gain and a qualitative one based on examining the changes in the nature of the processed signal. This is achieved by means of the recently introduced 'delay vector variance' (DVV) method for phase space signal characterisation. A comprehensive analysis of this approach on both linear and nonlinear benchmark signals suggests that the PRNN not only outperforms a single recurrent neural network (RNN) in terms of the prediction gain, but also has better or similar performance in terms of preserving the nature of the processed signal.
Mo Chen 0006, Danilo P. Mandic, Temujin Gautama, Marc M. Van Hulle
ICASSP (5)2
2005 A class of gradient-adaptive step size algorithms for complex-valued nonlinear neural adaptive filters
abstract
A class of variable step-size algorithms for complex-valued nonlinear neural adaptive finite impulse response (FIR) filters realised as a dynamical perceptron is proposed. The adaptive step-size is updated using gradient descent to give variable step-size complex-valued nonlinear gradient descent (VSCNGD) algorithms. These algorithms are shown to be capable of tracking signals with rich and unknown dynamics, and exhibit faster convergence and smaller steady state error than the standard algorithms. Further, the analysis of stability and computational complexity is provided. Simulations in the prediction setting support the approach.
Su Lee Goh, Danilo P. Mandic
ICASSP (5)2
2005 Semi-blind source separation for convolutive mixtures based on frequency invariant transformation
abstract
A novel method for separation of a class of convolutive mixtures is proposed, in which the received sensor signals are first transformed into instantaneous mixtures and then standard blind source separation (BSS) algorithms for instantaneous mixtures are applied. Since partial information about the mixing mechanism is required in the design of the transformation, the proposed method is strictly speaking semi-blind. From the beamforming viewpoint, the proposed approach represents a blind broadband beamforming method. As the separation is performed in fullband and only one separation is needed, the permutation problem associated with the frequency-domain BSS is avoided and the separation can be easily implemented online. Simulation results verify the usefulness of the proposed method.
Wei Liu 0001, Danilo P. Mandic
ICASSP (5)2
2004 Quality assessment of hybrid nonlinear filters
abstract
Traditionally, research on adaptive signal processing has been conducted with the aim of designing adaptive filters with high performance in terms of some prescribed performance measure. However, little is known about how such filters influence the nature of the processed signal. Based upon some recently introduced results in dealing with nonlinearity within a signal in hand, we provide a critical assessment of the qualitative performance of common linear and nonlinear filters and their combinations. An insight into the performance of so called hybrid filters is provided, which is achieved for combinations of standard nonlinear (neural) and linear filters. It is shown that depending on the application, it is important not only to look for best filter performance in terms of some quantitative measure of the error but also for a filter that will not change the character of a signal. Simulation results support the analysis.
Mo Chen 0006, Danilo P. Mandic
ICASSP (5)2
2004 Heteroscedastic kernel ridge regression
Gavin C. Cawley, Nicola L. C. Talbot, Robert J. Foxall, Stephen R. Dorling, Danilo P. Mandic
Neurocomputing5
2004 A Complex-Valued RTRL Algorithm for Recurrent Neural Networks
abstract
A complex-valued real-time recurrent learning (CRTRL) algorithm for the class of nonlinear adaptive filters realized as fully connected recurrent neural networks is introduced. The proposed CRTRL is derived for a general complex activation function of a neuron, which makes it suitable for nonlinear adaptive filtering of complex-valued nonlinear and nonstationary signals and complex signals with strong component correlations. In addition, this algorithm is generic and represents a natural extension of the real-valued RTRL. Simulations on benchmark and real-world complex-valued signals support the approach.
Su Lee Goh, Danilo P. Mandic
Neural Comput.2
2004 A novel adaptive learning rate sequential blind source separation algorithm
Maria G. Jafari, Jonathon A. Chambers, Danilo P. Mandic
Signal Process.3
2004 A generalized normalized gradient descent algorithm
abstract
A generalized normalized gradient descent (GNGD) algorithm for linear finite-impulse response (FIR) adaptive filters is introduced. The GNGD represents an extension of the normalized least mean square (NLMS) algorithm by means of an additional gradient adaptive term in the denominator of the learning rate of NLMS. This way, GNGD adapts its learning rate according to the dynamics of the input signal, with the additional adaptive term compensating for the simplifications in the derivation of NLMS. The performance of GNGD is bounded from below by the performance of the NLMS, whereas it converges in environments where NLMS diverges. The GNGD is shown to be robust to significant variations of initial values of its parameters. Simulations in the prediction setting support the analysis.
Danilo P. Mandic
IEEE Signal Process. Lett.1
2003 Approximately unbiased estimation of conditional variance in heteroscedastic kernel ridge regression
Gavin C. Cawley, Nicola L. C. Talbot, Robert J. Foxall, Stephen R. Dorling, Danilo P. Mandic
ESANN5
2003 A gradient adaptive step size algorithm for IIR filters
abstract
The output error method, a fundamental technique for the updating of the coefficients of an adaptive IIR filter, is modified by introducing a time varying step size. The adaptation of this term is based on a gradient descent technique. This scheme can be considered as an extension of the algorithms presented in Benveniste et al. (1990) and Mathews et al. (1993) to IIR filters. The novel algorithm does not require any a priori knowledge of the statistical characteristics of the input signal and the unknown channel, since its step size converges automatically to its optimal value. This algorithm has the ability to converge in time-varying environments, which makes it suitable for processing of nonstationary signals.
Christos Boukis, Danilo P. Mandic, Eftychios V. Papoulis, Anthony G. Constantinides
ICASSP (6)2
2003 A differential entropy based method for determining the optimal embedding parameters of a signal
abstract
A novel method for determining the set of parameters for a phase space representation of a time series is proposed. Based upon the differential entropy, both the optimal embedding dimension m, and time lag /spl tau/, are simultaneously determined. The choice of these parameters is closely related to the length of the optimal tap input delay line of an adaptive filter or time-delay neural network. The method employs a single criterion - the "entropy ratio" between the phase space representation of a signal and an ensemble of its surrogates - and is first systematically tested on synthetic time series for which the optimal embedding parameters are known, after which it is verified on a number of benchmark real-world time series. The proposed entropy ratio method is shown to consistently outperform some well-established methods.
Temujin Gautama, Danilo P. Mandic, Marc M. Van Hulle
ICASSP (6)2
2003 Analysis of the class of complex-valued error adaptive normalised nonlinear gradient descent algorithms
abstract
A complex-valued gradient based algorithm for training nonlinear complex-valued finite impulse response (FIR) filters is derived. The proposed complex error-adaptive normalised nonlinear gradient descent (CEANNGD) and smoothed CEANNGD (SCEANNGD) algorithms are an improvement on the complex nonlinear gradient descent (CNGD) and the complex normalised nonlinear gradient descent (CNNGD) algorithm by including an adaptive term in the normalised learning rate of the CNNGD. This is achieved by performing a minimisation of the complex-valued instantaneous output error that has been approximated via a Taylor series expansion, which makes it suitable for the processing of nonlinear and nonstationary signals. Experiments on complex-valued coloured and nonlinear signals show that the CEANNGD and SCEANNGD algorithms outperform the standard CNNGD and CNGD algorithms.
Andrew I. Hanna, Ian Yates, Danilo P. Mandic
ICASSP (2)3
2003 A normalized mixed-norm adaptive filtering algorithm robust under impulsive noise interference
abstract
A normalized robust mixed-norm (NRMN) algorithm for system identification in the presence of impulsive noise is introduced. The standard robust mixed-norm (RMN) algorithm, despite its ability to cope with impulsive noise by virtue of combining the first and second error norm in the cost function it minimizes, exhibits slow convergence, requires a stationary operating environment, and employs a constant step-size which needs to be determined a-priori. To overcome these limitations, the proposed NRMN algorithm introduces a time varying learning rate which is derived based upon the dynamics of the input signal, and thus no longer requires a stationary environment, a major drawback of the RMN algorithm. The normalized step-size is bounded from above and a parameter is introduced within its upper-bound, which provides a trade-off between the convergence rate and the steady-state coefficient error. The analysis and experimental results show that the proposed NRMN exhibits increased convergence rate and substantially reduces the steady-state coefficient error, as compared to the least absolute deviation (LAD) and RMN algorithms.
Danilo P. Mandic, Eftychios V. Papoulis, Christos Boukis
ICASSP (6)1
2003 A Non-parametric Test for Detecting the Complex-Valued Nature of Time Series
Temujin Gautama, Danilo P. Mandic, Marc M. Van Hulle
KES2
2003 A Data-Reusing Gradient Descent Algorithm for Complex-Valued Recurrent Neural Networks
Su Lee Goh, Danilo P. Mandic
KES2
2003 A Fast Converging Sequential Blind Source Separation Algorithm for Cyclostationary Sources
Maria G. Jafari, Danilo P. Mandic, Jonathon A. Chambers
KES2
2003 Recurrent neural networks with trainable amplitude of activation functions
Su Lee Goh, Danilo P. Mandic
Neural Networks2
2003 A complex-valued nonlinear neural adaptive filter with a gradient adaptive amplitude of the activation function
Andrew I. Hanna, Danilo P. Mandic
Neural Networks2
2003 A Data Reusing Nonlinear Gradient Descent Algorithm for a Class of Complex Valued Neural Adaptive Filters
Andrew I. Hanna, Danilo P. Mandic
Neural Process. Lett.2
2003 Signal Nonlinearity in fMRI: A Comparison between BOLD and MION
abstract
In this paper, we introduce a methodology for comparing the nonlinearities present in sets of time series using four different nonlinearity measures, one of which, the "delay vector variance" method, is a novel approach to the characterization of a time series. It is then applied to examine the difference in nonlinearity between functional magnetic resonance imaging (fMRI) signals that have been recorded using different contrast agents. Recently, an exogenous contrast agent, monocrystalline iron oxide particle (MION), has been introduced for fMRI, which has been shown to increase the functional sensitivity compared with the traditional blood oxygen level dependent (BOLD) technique. The resulting fMRI signals are influenced by cerebral blood volume, whereas the more traditionally recorded BOLD signals are influenced not only by cerebral blood volume, but also by the cerebral blood flow and the metabolic rate of oxygen. The proposed methodology is applied to address the question whether this difference in the number of physiological variables is reflected in a difference in the degree of nonlinearity. We therefore analyze two sets of fMRI signals, one from a BOLD and the other from a MION monkey study with similar experimental designs. In the neuroimaging context, the proposed nonlinearity analyses are different from those described in the literature, since no a priori model is assumed: rather than pinpointing the source(s) of nonlinearity, nonparametric analyses are performed on BOLD and MION fMRI signals. Furthermore, we introduce a strategy for analyzing a population of fMRI signals, rather than focusing the analysis on one signal, as is traditionally done in the domain of nonlinear signal processing. Our results show that, overall, the BOLD signals are more nonlinear in nature than the MION ones, which is in agreement with current hypotheses.
Temujin Gautama, Danilo P. Mandic, Marc M. Van Hulle
IEEE Trans. Medical Imaging2
2002 Heteroscedastic regularised kernel regression for prediction of episodes of poor air quality
Robert J. Foxall, Gavin C. Cawley, Nicola L. C. Talbot, Stephen R. Dorling, Danilo P. Mandic
ESANN5
2002 Error Functions for Prediction of Episodes of Poor Air Quality
Robert J. Foxall, Gavin C. Cawley, Stephen R. Dorling, Danilo P. Mandic
ICANN4
2002 A normalised complex backpropagation algorithm
abstract
A backpropagation based algorithm for training nonlinear complex valued feed-forward neural networks employed as nonlinear adaptive filters is derived. The proposed normalised complex backpropagation (NCBP) algorithm is an improvement on the complex backpropagation (CBP) algorithm by including an adaptive normalised learning rate. This is achieved by performing a minisation of the complex-valued instantaneous output error that has been expanded via a Taylor series expansion. The proposed algorithm is applicable to any complex-valued nonlinear architecture. Experiments on complex coloured and nonlinear signals confirm that the NCBP algorithm outperforms the standard CBP algorithm.
Andrew I. Hanna, Danilo P. Mandic
ICASSP2
2002 On the derivation of the optimal payload size for packet based transmission over a binary symmetrical communication channel
abstract
A criterion function for selecting the optimal payload size in data transmission over a binary symmetrical communication channel is presented. The work assumes that the packets comprise a fixed size header and variable length payload. The conventional criterion function is based on the ratio between payload size and mean number of transmitted bits per packet, which includes re-transmissions of corrupted packets. Applying a logarithm to the inverse of this criterion function makes the resulting function very suitable for mathematical analysis. This also facilitates parameter sensitivity studies without recourse to numerical methods. As such, this helps to visualise and understand the problem of optimal payload length selection when teaching courses in signal processing for communications and multimedia.
Djemal H. Kolonic, Danilo P. Mandic, Ben P. Milner, Richard W. Harvey
ICASSP2
2002 A general adaptive normalised nonlinear-gradient descent algorithm for nonlinear adaptive filters
abstract
An algorithm for training nonlinear adaptive finite impulse response (FIR) filters employed for nonlinear prediction and system identification is introduced. This general adaptive normalised nonlinear gradient descent (ANNGD) algorithm is fully gradient adaptive, unlike previously proposed algorithms of this kind. It is derived based upon the Taylor series expansion of the instantaneous output error of the filter. For rigour, the remainder of the Taylor series expansion in the derivation of the algorithm is made adaptive thus providing an adaptive learning rate. Experiments on coloured and nonlinear signals confirm that the ANNGD outperforms the other algorithms of this kind.
Danilo P. Mandic, Andrew I. Hanna, Dai I. Kim
ICASSP1
2002 On sensitivity of neural adaptive filters with respect to the slope parameter of a neuron activation function
abstract
Sensitivity analysis of neural adaptive filters with respect to the slope parameter of a neuron activation function is performed. The analysis is provided both for a feedforward neural adaptive filter and a recurrent perceptron. The slope affects stability and convergence characteristics of a filter via inherent relationship between the slope and the learning rate parameter. In addition, it determines character of an activation function, i.e. whether it is contractive or expansive mapping. Presented analysis shows that gradient-descent based learning algorithms with an adaptive learning rate significantly reduce sensitivity of a neural adaptive filter with respect to the slope parameter, when compared with learning algorithms with a constant learning rate. Experimental results on the test speech and HRV signals support the analysis.
Warren Sherliker, Igor R. Krcmar, Milorad M. Bozic, Danilo P. Mandic
ICASSP4
2002 Data-Reusing Recurrent Neural Adaptive Filters
abstract
A class of data-reusing learning algorithms for real-time recurrent neural networks (RNNs) is analyzed. The analysis is undertaken for a general sigmoid nonlinear activation function of a neuron for the real time recurrent learning training algorithm. Error bounds and convergence conditions for such data-reusing algorithms are provided for both contractive and expansive activation functions. The analysis is undertaken for various configurations that are generalizations of a linear structure infinite impulse response adaptive filter.
Danilo P. Mandic
Neural Comput.1
2002 Nonlinear FIR adaptive filters with a gradient adaptive amplitude in the nonlinearity
abstract
A nonlinear gradient descent (NGD) learning algorithm with an adaptive amplitude of the nonlinearity is derived for the class of nonlinear finite impulse response (FIR) adaptive filters (dynamical perceptron). This is based on the adaptive amplitude backpropagation (AABP) algorithm for large-scale neural networks. The amplitude of the nonlinear activation function is made gradient adaptive to give the adaptive amplitude nonlinear gradient descent (AANGD) algorithm, making the AANGD suitable for processing nonlinear and nonstationary input signals with a large dynamical range. Experimental results show the AANGD algorithm outperforming the standard NGD algorithm on both colored and nonlinear input with large dynamics. Despite its simplicity, the considered algorithm proves suitable for adaptive filtering of nonlinear and nonstationary signals.
Andrew I. Hanna, Danilo P. Mandic
IEEE Signal Process. Lett.2
2001 Nonlinear modelling of air pollution time series
abstract
An analysis of predictability of a nonlinear and nonstationary ozone time series is provided. For rigour, the deterministic versus stochastic (DVS) analysis is first undertaken to detect and measure inherent nonlinearity of the data. Based upon this, neural and linear adaptive predictors are compared on this time series for various filter orders, hence indicating the embedding dimension. Simulation results confirm the analysis and show that for this class of air pollution data, neural, especially recurrent neural predictors, perform best.
Robert J. Foxall, Igor R. Krcmar, Gavin C. Cawley, Stephen R. Dorling, Danilo P. Mandic
ICASSP5
2001 A fully adaptive normalized nonlinear gradient descent algorithm for nonlinear system identification
abstract
A fully adaptive normalized nonlinear gradient descent (FANNGD) algorithm for neural adaptive filters employed for nonlinear system identification is proposed. This full adaptation is achieved using the instantaneous squared prediction error to adapt the free parameter of the NNGD algorithm. The convergence analysis of the proposed algorithm is undertaken using the contractivity property of the nonlinear activation function of a neuron. Simulation results show that a fully adaptive NNGD algorithm outperforms the standard NNGD algorithm for nonlinear system identification.
Igor R. Krcmar, Danilo P. Mandic
ICASSP2
2001 A normalized gradient descent algorithm for nonlinear adaptive filters using a gradient adaptive step size
abstract
A fully adaptive normalized nonlinear gradient descent (FANNGD) algorithm for online adaptation of nonlinear neural filters is proposed. An adaptive stepsize that minimizes the instantaneous output error of the filter is derived using a linearization performed by a Taylor series expansion of the output error. For rigor, the remainder of the truncated Taylor series expansion within the expression for the adaptive learning rate is made adaptive and is updated using gradient descent. The FANNGD algorithm is shown to converge faster than previously introduced algorithms of this kind.
Danilo P. Mandic, Andrew I. Hanna, Moe Razaz
IEEE Signal Process. Lett.1
2000 A normalized gradient algorithm for an adaptive recurrent perceptron
abstract
A normalized algorithm for on-line adaptation of a recurrent perceptron is derived. The algorithm builds upon the normalized backpropagation (NBP) algorithm for feedforward neural networks, and provides an adaptive learning rate and normalization for a recurrent perceptron learning algorithm. The algorithm is based upon local linearization about the current point in the state-space of the network. Such a learning rate is normalized by the squared norm of the gradient at the neuron, which extends the notion of normalized linear algorithms to the nonlinear case.
Jonathon A. Chambers, Warren Sherliker, Danilo P. Mandic
ICASSP3
2000 Visualising error surfaces for adaptive filters and other purposes
abstract
Modern neural and adaptive systems often have complicated error performance surfaces with many local extrema. Visualising and understanding these surfaces is critical to effective tuning of these systems but almost all visualisation methods are confined to two dimensions. Here we show how to use a morphological scale-space transform to convert these multi-dimensional complex error surfaces into two-dimensional trees where the leaf nodes are local minima and other nodes represent decision points such as saddle points and points of inflection.
Mark Fisher 0001, Danilo P. Mandic, J. Andrew Bangham, Richard W. Harvey
ICASSP2
2000 Some potential pitfalls with s to z-plane mappings
abstract
Design of digital infinite impulse response (IIR) filters is a compulsory topic in most signal processing courses. Most often, it is taught by using the bilinear transform to map an analogue counterpart into the corresponding digital filter. The usual approach is to define a mapping between the complex variables s and z, and hence, by substitution, derive a mapping between /spl omega/, analogue frequency, and /spl theta/, sampled frequency. This is rather elliptical, since the real aim is to establish the correspondence between the frequency response of a prototype analogue system H(j/spl omega/), and H(e/sup j/spl theta//), the response of the sampled system. Here we provide a rigorous analysis for the mutual invertibility between the analogue frequency /spl omega/, and the digital frequency /spl theta/ for this case. Based upon the definition of the tan and arctan functions, conditions of existence, uniqueness and continuity of such a mutually inverse mapping are derived. Based upon these results, simple proofs for the mutually inverse mappings /spl omega//spl rarr//spl theta/ and /spl theta//spl rarr//spl omega/ are given. This is supported by appropriate diagrams. This problem arose as a student question while teaching DSP.
Richard W. Harvey, Danilo P. Mandic, Djemal H. Kolonic
ICASSP2
2000 On global asymptotic stability of fully connected recurrent neural networks
abstract
Conditions for global asymptotic stability (GAS) of a nonlinear relaxation process realized by a recurrent neural network (RNN) are provided. Existence, convergence, and robustness of such a process are analyzed. This is undertaken based upon the contraction mapping theorem (CMT) and the corresponding fixed point iteration (FPI). Upper bounds for such a process are shown to be the conditions of convergence for a commonly analyzed RNN with a linear state dependence.
Danilo P. Mandic, Jonathon A. Chambers, Milorad M. Bozic
ICASSP1
2000 Relationships Between the A Priori and A Posteriori Errors in Nonlinear Adaptive Neural Filters
abstract
The lower bounds for the a posteriori prediction error of a nonlinear predictor realized as a neural network are provided. These are obtained for a priori adaptation and a posteriori error networks with sigmoid nonlinearities trained by gradient-descent learning algorithms. A contractivity condition is imposed on a nonlinear activation function of a neuron so that the a posteriori prediction error is smaller in magnitude than the corresponding a priori one. Furthermore, an upper bound is imposed on the learning rate eta so that the approach is feasible. The analysis is undertaken for both feedforward and recurrent nonlinear predictors realized as neural networks.
Danilo P. Mandic, Jonathon A. Chambers
Neural Comput.1
2000 Towards the Optimal Learning Rate for Backpropagation
Danilo P. Mandic, Jonathon A. Chambers
Neural Process. Lett.1
2000 A normalised real time recurrent learning algorithm
Danilo P. Mandic, Jonathon A. Chambers
Signal Process.1
2000 On the choice of parameters of the cost function in nested modular RNN's
abstract
We address the choice of the coefficients in the cost function of a modular nested recurrent neural-network (RNN) architecture, known as the pipelined recurrent neural network (PRNN). Such a network can cope with the problem of vanishing gradient, experienced in prediction with RNN's. Constraints on the coefficients of the cost function, in the form of a vector norm, are considered. Unlike the previous cost function for the PRNN, which included a forgetting factor motivated by the recursive least squares (RLS) strategy, the proposed forms of cost function provide "forgetting" of the outputs of adjacent modules based upon the network architecture. Such an approach takes into account the number of modules in the PRNN, through the unit norm constraint on the coefficients of the cost function of the PRNN. This is shown to be particularly suitable, since due to inherent nesting in the PRNN, every module gives its full contribution to the learning process, whereas the unit norm constrained cost function introduces a sense of forgetting in the memory management of the PRNN. The PRNN based upon a modified cost function outperforms existing PRNN schemes in the time series prediction simulations presented.
Danilo P. Mandic, Jonathon A. Chambers
IEEE Trans. Neural Networks Learn. Syst.1
1999 Global asymptotic convergence of nonlinear relaxation equations realised through a recurrent perceptron
abstract
Conditions for global asymptotic stability (GAS) of a nonlinear relaxation equation realised by a nonlinear autoregressive moving average (NARMA) recurrent perceptron are provided. Convergence is derived through fixed point iteration (FPI) techniques, based upon a contraction mapping feature of a nonlinear activation function of a neuron. Furthermore, nesting is shown to be a spatial interpretation of an FPI, which underpins a pipelined recurrent neural network (PRNN) for nonlinear signal processing.
Danilo P. Mandic, Jonathon A. Chambers
ICASSP1
1999 Relating the Slope of the Activation Function and the Learning Rate Within a Recurrent Neural Network
abstract
A relationship between the learning rate η in the learning algorithm, and the slope β in the nonlinear activation function, for a class of recurrent neural networks (RNNs) trained by the real-time recurrent learning algorithm is provided. It is shown that an arbitrary RNN can be obtained via the referent RNN, with some deterministic rules imposed on its weights and the learning rate. Such relationships reduce the number of degrees of freedom when solving the nonlinear optimization task of finding the optimal RNN parameters.
Danilo P. Mandic, Jonathon A. Chambers
Neural Comput.1
1999 Exploiting inherent relationships in RNN architectures
Danilo P. Mandic, Jonathon A. Chambers
Neural Networks1
1999 Toward an optimal PRNN-based nonlinear predictor
abstract
We present an approach for selecting optimal parameters for the pipelined recurrent neural network (PRNN) in the paradigm of nonlinear and nonstationary signal prediction. Although there has recently been progress in terms of algorithms for training the PRNN, no account has been made of some inherent features of the PRNN. We therefore provide a study of the role of nesting, which is inherent to the PRNN architecture. The corresponding number of nested modules needed for a certain prediction task, and their contribution toward the final prediction gain (PG) give a thorough insight into the way the PRNN performs, and offers solutions for optimization of its parameters. In particular, nesting, which is a contractive function by its nature, allows the forgetting factor in the cost function of the PRNN to exceed unity, hence it becomes an emphasis factor. This compensates for the small contribution of the distant modules to the prediction process, due to nesting, and helps to circumvent the problem of vanishing gradient, experienced in RNN's for prediction. The PRNN, with its parameters chosen based upon the established criteria, is shown to outperform the linear least mean square (LMS) and recursive least squares (RLS) predictors, as well as previously proposed PRNN schemes, at no expense of additional computational complexity.
Danilo P. Mandic, Jonathon A. Chambers
IEEE Trans. Neural Networks1