Zongben Xu

dblp:25/3264 · also Zong-Ben Xu · DBLP profile ↗
← Back
290ranked-venue papers
19as first author
87since 2021 · last 2026
0000-0002-4066-2338ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 162 · 10 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 70 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 1 first-author · 8 since 2021Computer networks · 16 · 13 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorTheory of computation · 3
YearPublicationVenuePosition
2026 G-IR: Geometric Image Representation for Learning
abstract
Images are generally represented by pixel intensities or color values, which are usually used as direct inputs for learning. This study innovatively proposes a geometric image representation method and refreshes the general learning model (e.g., autoencoder) in the diffeomorphic space. Based on the theory of geometric optimal transport and quasiconformal mapping, we equivalently transform the intensity representation into a shape representation. The image space becomes a diffeomorphic space, where any image can be uniquely represented as a Beltrami coefficient function defined on a uniform grid reference, and vice versa. This innovative geometric image representation (G-IR) captures the fine-grained structure inherent in the entire image, which is different from the traditional feature extraction that focuses on the internal geometric objects of the image (such as boundaries and axes). The diffeomorphic property preserves structure in the generation process, which is very necessary in the field of real physics. It can be assembled into existing pipelines as a plug-in, providing structure-preserving properties for the entire framework. Experiments on image restoration and interpolation validated the high efficiency, efficacy and applicability of the G-IR method, demonstrating its superior performance compared to common pixel-level image appearance representations.
Zongben Xu
AAAI4
2026 DAC-MR: Data Augmentation Consistency Based Meta-Regularization for Meta-Learning
abstract
Meta learning recently has been heavily researched and helped advance the contemporary machine learning. However, achieving well-performing meta-learning model requires a large amount of training tasks with high-quality meta-data representing the underlying task generalization goal, which is sometimes difficult and expensive to obtain for real applications. Current meta-data-driven meta-learning approaches, however, are fairly hard to train satisfactory meta-models with imperfect training tasks. To address this issue, we suggest a meta-knowledge informed meta-learning (MKIML) framework to improve meta-learning by additionally integrating compensated meta-knowledge into meta-learning process. We preliminarily integrate task-agnostic meta-knowledge into meta-objective via using an appropriate meta-regularization (MR) objective to regularize capacity complexity of the meta-model function class to facilitate better generalization on unseen tasks. As a practical implementation, we introduce data augmentation consistency to encode invariance as meta-knowledge for instantiating MR objective, denoted by DAC-MR. The proposed DAC-MR is hopeful to learn well-performing meta-models from training tasks with noisy, sparse or unavailable meta-data. We theoretically demonstrate that DAC-MR can be treated as a proxy meta-objective used to evaluate meta-model without high-quality meta-data. Besides, meta-data-driven meta-loss objective combined with DAC-MR is capable of achieving better meta-level generalization. 12 meta-learning tasks with different network architectures and benchmarks substantiate the capability of our DAC-MR on aiding meta-model learning. Fine performance of DAC-MR are obtained across all settings, and are well-aligned with our theoretical insights. This implies that our DAC-MR is problem-agnostic, and hopeful to be readily applied to extensive meta-learning problems and tasks. All codes for reproducing our experimental results are released athttps://github.com/xjtushujun/DAC-MR.
Deyu Meng, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Scalable Pre-Trained Masked Channel Model of Wireless Communications
abstract
Deep learning (DL)-based models have been widely applied in wireless communication systems with excellent performance. However, most of these models are task- and scenario-specific, exhibiting limited generalization and contributing to increasing complexity and overhead with their deployment in systems. Inspired by the emergent capabilities and strong generalization exhibited by large models (LMs), represented by large language models (LLMs), this paper analyzes the differences between existing DL-based wireless communication models and LLMs, proposing a framework for designing LMs tailored to wireless communications. Building upon this framework, we integrate channel-related tasks of the physical layer into a unified pre-training task, i.e., channel completion, and propose a pre-trained masked channel model (MCM) with different parameter scales ranging from 5 million to 1 billion (B), enabling simultaneous solving of channel state information (CSI) feedback, prediction, and estimation. Additionally, scaling laws on these downstream tasks are derived to guide the design and deployment of MCMs. The formulated scaling laws indicate that the proposed MCM with 1B parameter not only shows no sign of performance saturation on the pre-trained task but also has the potential to enhance performance at larger model sizes. Simulation results demonstrate that the proposed MCM outperforms the existing algorithms across various downstream tasks while exhibiting superior cross-task and cross-scenario generalization capabilities in both simulated and realistic scenarios.
Zhongsheng Deng, Zhen Qiao, Jiang Xue 0001, Dusit Niyato, Zongben Xu
IEEE Trans. Commun.7
2026 Neuro-PLS: A Generalizable Local Search Framework for Multiobjective Combinatorial Optimization
Haotian Zhang 0023, Jialong Shi, Jianyong Sun, Qingfu Zhang 0001, Zongben Xu
IEEE Trans. Evol. Comput.5
2026 Design Intelligent Air Interface of MIMO Systems
abstract
The architectural design of the air interface plays a critical role in wireless communications, embedding crucial functionality to guarantee both efficiency and robustness. Physical layer algorithms often face performance challenges in real-world scenarios owing to the basic assumptions of Gaussian noise, channel model linearity, and functional separation. This paper explores intelligent air interface (IAI) algorithms for multiple-input multiple-output (MIMO) systems to overcome the limitations of these assumptions. The physical layer link is restructured as a composite of various functions and framed as a mathematical optimization problem aimed at maximizing transmission rates, solved through optimization sub-problems for each function using specialized neural networks. Additionally, this paper presents the intelligent modulation and demodulation network (IMD-Net) with an adaptive adjustment sub-network, joint channel feedback and prediction network (CFP-Net), and GEM-Net for joint channel estimation and signal detection, using an unfolded generalized expectation maximization algorithm. Simulation results indicate that the proposed algorithms surpass traditional linear methods and the independent deep learning (DL) based methods in various scenarios and configurations.
Runhua Li, Guanzhang Liu, Zhengyang Hu 0001, Yiqing Zhang 0001, Feng Li 0057, Jiang Xue 0001, John S. Thompson, Zongben Xu
IEEE Trans. Wirel. Commun.10
2026 Intelligent Predictive Beamforming for Integrated Sensing, Communication and Power Transfer for Low-Altitude Economy
abstract
This paper investigates intelligent predictive beamforming design for simultaneous wireless information and power transfer-integrated sensing and communication (SWIPT-ISAC) systems for low-altitude economy wireless networks. Considering the downlink scenario where the base station aims to localize the moving targets/communication users and also transfer power to them, we formulate a weighted sum optimization problem to balance the trade-off between achievable communication rate and harvested energy, subject to sensing accuracy constraints defined by the Cramér–Rao lower bound. To address the non-convexity of the problem, we propose the Time-Spatial Fusion Network (TSFusionNet), an unsupervised deep learning (DL) framework that leverages multi-step historical channel state information for predictive beamforming design. TSFusionNet integrates convolutional and recurrent layers with a differential attention mechanism to capture spatial-temporal dependencies and mitigate non-stationary channel dynamics. We introduce a dynamic penalty-based loss function to enforce sensing constraints during training. Simulation results show that by adjusting the weight factor, the proposed method achieves a trade-off between rate and energy while meeting sensing accuracy requirements. Moreover, it significantly reduces computational complexity by up to approximately 96.8% in parameters and 81.5% in FLOPs, compared to existing DL frameworks.
Faheem Ahmad Khan, Zhiqiang Wei 0001, Jiang Xue 0001, Christos Masouros, Dusit Niyato, Zongben Xu
IEEE Trans. Wirel. Commun.7
2025 Improving Memory Efficiency for Training KANs via Meta Learning
abstract
Inspired by the Kolmogorov-Arnold representation theorem, KANs offer a novel framework for function approximation by replacing traditional neural network weights with learnable univariate functions. This design demonstrates significant potential as an efficient and interpretable alternative to traditional MLPs. However, KANs are characterized by a substantially larger number of trainable parameters, leading to challenges in memory efficiency and higher training costs compared to MLPs. To address this limitation, we propose to generate weights for KANs via a smaller meta-learner, called MetaKANs. By training KANs and MetaKANs in an end-to-end differentiable manner, MetaKANs achieve comparable or even superior performance while significantly reducing the number of trainable parameters and maintaining promising interpretability. Extensive experiments on diverse benchmark tasks, including symbolic regression, partial differential equation solving, and image classification, demonstrate the effectiveness of MetaKANs in improving parameter efficiency and memory usage. The proposed method provides an alternative technique for training KANs, that allows for greater scalability and extensibility, and narrows the training cost gap with MLPs stated in the original paper of KANs. Our code is available at https://github.com/Murphyzc/MetaKAN.
Zhangchi Zhao, Deyu Meng, Zongben Xu
ICML4
2025 Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport
abstract
Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i.e., retrospective reconstruction. However, training on simulated pairs commonly leads to performance degradation on real prospective data due to the retrospective-to-prospective gap caused by incomplete imaging knowledge in simulation. To address this challenge, this paper introduces imaging Knowledge-Informed Dynamic Optimal Transport (KIDOT), a novel dynamic optimal transport framework with optimality in the sense of preserving consistency with imaging physics in transport, that conceptualizes reconstruction as finding a dynamic transport path. KIDOT learns from unpaired data by modeling reconstruction as a continuous evolution path from measurements to images, guided by an imaging knowledge-informed cost function and transport equation. This dynamic and knowledge-aware approach enhances robustness and better leverages unpaired data while respecting acquisition physics. Theoretically, we demonstrate that KIDOT naturally generalizes dynamic optimal transport, ensuring its mathematical rationale and solution existence. Extensive experiments on MRI and CT reconstruction demonstrate KIDOT's superior performance. Code is available at https://github.com/TaoranZheng717/KIDOT.
Taoran Zheng, Yan Yang 0007, Xing Li 0027, Xiang Gu 0005, Jian Sun 0009, Zongben Xu
NeurIPS6
2025 Self-supervised distributional and contrastive learning model for image anomaly detection
Yannan Pu, Jian Sun 0009, Niansheng Tang, Zongben Xu
Knowl. Based Syst.4
2025 Training Networks in Null Space of Feature Covariance With Self-Supervision for Incremental Learning
abstract
In the context of incremental learning, a network is sequentially trained on a stream of tasks, where data from previous tasks are particularly assumed to be inaccessible. The major challenge is how to overcome the stability-plasticity dilemma, i.e., learning knowledge from new tasks without forgetting the knowledge of previous tasks. To this end, we propose two mathematical conditions for guaranteeing network stability and plasticity with theoretical analysis. The conditions demonstrate that we can restrict the parameter update in the null space of uncentered feature covariance at each linear layer to overcome the stability-plasticity dilemma, which can be realized by layerwise projecting gradient into the null space. Inspired by it, we develop two algorithms, dubbed Adam-NSCL and Adam-SFCL respectively, for incremental learning. Adam-NSCL and Adam-SFCL provide different ways to compute the projection matrix. The projection matrix in Adam-NSCL is constructed by singular vectors associated with the smallest singular values of the uncentered feature covariance matrix, while the projection matrix in Adam-SFCL is constructed by all singular vectors associated with adaptive scaling factors. Additionally, we explore adopting self-supervised techniques, including self-supervised label augmentation and a newly proposed contrastive loss, to improve the performance of incremental learning. These self-supervised techniques are orthogonal to Adam-NSCL and Adam-SFCL and can be incorporated with them seamlessly, leading to Adam-NSCL-SSL and Adam-SFCL-SSL respectively. The proposed algorithms are applied to task-incremental and class-incremental learning on various benchmark datasets with multiple backbones, and the results show that they outperform the compared incremental learning methods.
Shipeng Wang 0002, Xiaorong Li, Jian Sun 0009, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Rotation Equivariant Arbitrary-Scale Image Super-Resolution
abstract
The arbitrary-scale image super-resolution (ASISR), a recent popular topic in computer vision, aims to achieve arbitrary-scale high-resolution recoveries from a low-resolution input image. This task is realized by representing the image as a continuous implicit function through two fundamental modules, a deep-network-based encoder and an implicit neural representation (INR) module. Despite achieving notable progress, a crucial challenge of such a highly ill-posed setting is that many common geometric patterns, such as repetitive textures, edges, or shapes, are seriously warped and deformed in the low-resolution images, naturally leading to unexpected artifacts appearing in their high-resolution recoveries. Embedding rotation equivariance into the ASISR network is thus necessary, as it has been widely demonstrated that this enhancement enables the recovery to faithfully maintain the original orientations and structural integrity of geometric patterns underlying the input image. Motivated by this, we make efforts to construct a rotation equivariant ASISR method in this study. Specifically, we elaborately redesign the basic architectures of INR and encoder modules, incorporating intrinsic rotation equivariance capabilities beyond those of conventional ASISR networks. Through such amelioration, the ASISR network can, for the first time, be implemented with end-to-end rotational equivariance maintained from input to output. We also provide a solid theoretical analysis to evaluate its intrinsic equivariance error, demonstrating its inherent nature of embedding such an equivariance structure. The superiority of the proposed method is substantiated by experiments conducted on both simulated and real datasets. We also validate that the proposed framework can be readily integrated into current ASISR methods in a plug & play manner to further enhance their performance.
Qi Xie 0002, Jiahong Fu, Zongben Xu, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Adaptive representation learning and sample weighting for low-quality 3D face recognition
Cuican Yu, Fengxun Sun, Huibin Li 0001, Liming Chen 0002, Jian Sun 0009, Zongben Xu
Pattern Recognit.7
2025 Protein Structure Prediction Using a New Optimization-Based Evolutionary and Explainable Artificial Intelligence Approach
abstract
Protein structure prediction (PSP) is an important scientific problem because it helps humans to understand how proteins perform their biological functions. This paper models the PSP problem as a multi-objective optimization problem with three fast and accurate knowledge-based energy functions. This way, using evolutionary computation (EC)-based artificial intelligence (AI) approach to solve this multi-objective PSP problem to find the optimal structure is explainable. Considering that the multiple populations for multiple objectives (MPMO) framework shows efficient performance in solving lots of multi-objective benchmarks and real-world problems, this paper proposes a new AI approach named improved MPMO-based differential evolution (IMPMO-DE) to solve the multi-objective PSP problem. To our best knowledge, this is the first time that MPMO is applied to PSP, with three novel strategies. First, an adaptive archive-based mutation strategy is proposed to better balance the exploration and exploitation abilities by adaptively using different archive-based mutation operators in different evolutionary stages. Second, a mixed individual transfer strategy is proposed to share search information among the multiple populations to accelerate the convergence speed. Third, an evolvable archive update strategy is proposed to generate more promising solutions through evolving the archived solutions. IMPMO-DE is tested on 28 representative proteins and all the available template-free modeling proteins up to 404 residues in the famous Critical Assessment of Protein Structure Prediction (CASP14) competition. Experimental results show that IMPMO-DE performs better than the compared state-of-the-art EC-based PSP methods and ranks above average compared with all the CASP14 competitors. More importantly, IMPMO-DE is a new efficient AI approach that opens a promising optimization-based evolutionary and explainable way for efficient PSP rather than deep learning approaches like AlphaFold2, especially for newly discovered proteins without similar known protein structures.
Zhi-hui Zhan, Langchong He, Zongben Xu, Jun Zhang 0003
IEEE Trans. Evol. Comput.4
2025 LSTM Network Assisted Construction of the Angle-Dependent Point Spread Function and Its Applications in Seismic Imaging
abstract
Migration is the core link in reflection seismic exploration. Seismic images are often extended into angle domain for interpretations. However, affected by limited acquisition aperture and complex overburden, the generated images are far from ideal. The illumination is unbalanced, causing unreliable amplitude variation in angle gathers. Band-limited seismic data and wavelet stretch in large angles, lead to low-resolution angle gathers. Image-domain least-squares migration (IDLSM) implemented by point spread function (PSF) deconvolution is a promising solution. Extending the concept of IDLSM to the angle domain, we develop a new method to construct angle-dependent PSFs and optimize angle gathers. The essential element to construct PSFs is the Green’s function. The proposed method reconstructs Green’s functions using a bidirectional long short-term memory (LSTM) network. We use a ray tracing method to efficiently obtain wave propagation directions (travel-time gradients). And wave-equation forward modeling is used to accurately calculate wavefront amplitudes. The LSTM network is trained by labels composed of travel-time gradients and amplitudes to surrogate the solver of Green’s functions. Angle-dependent PSFs are constructed according to the mathematical model of the angular local Hessian. And inversions with PSFs are performed to optimize angle gathers. Numerical tests on a 3-D synthetic model demonstrate that the proposed method is able to improve the image quality of prestack angle gathers and poststack seismic images. The proposed method compensates illumination and improves the resolution of angle-dependent seismic images. Both vertical and lateral resolution are enhanced. Amplitude versus angle (AVA) responses can be corrected for further analysis and interpretations.
Feipeng Li, Jinghuai Gao, Zhiguo Wang 0002, Chuang Li 0003, Zhaoqi Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.7
2025 True Amplitude Seismic Imaging With Wave Equation-Based Illumination Compensation in the Dip and Reflection Angle Domain
abstract
Seismic interpretation and reservoir characterization require the seismic data having faithful amplitudes that relate to subsurface physical parameters. Nowadays, the amplitude fidelity of seismic imaging becomes more important than ever. Although reverse time migration (RTM) adopts the full wave equation as true amplitude seismic wave propagator, it is still not sufficient for true amplitude seismic imaging since migration is only the adjoint operator corresponding to the forward modeling process. The complex overburden and limited migration aperture lead to unbalanced illumination of subsurface structures. Least-squares migration was proposed to correct amplitudes of seismic images, but it is computationally expensive and sometimes unstable. The illumination compensation is an available alternative which only considers the amplitude correction regardless of the resolution issue. In this article, we propose a true amplitude seismic imaging method with illumination compensation performed on both RTM stacked images and angle gathers. We derive the angle-dependent illumination intensity from the Hessian of least-squares migration in which Green’s functions are essential components. We propose a new method to estimate the Green’s function and its corresponding wave propagation direction based on wavefields excitation amplitudes and Poynting vectors at excitation times. Then, the illumination intensity is constructed as a function of dip and reflection angles to correct both angle gathers and stacked images. The proposed method is tested using two synthetic models and a real marine dataset. Numerical results demonstrate that the proposed method can effectively correct amplitudes of seismic images. Deep events beneath complex structures are enhanced with more balance illumination.
Feipeng Li, Jinghuai Gao, Zhiguo Wang 0002, Chuang Li 0003, Zhaoqi Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.6
2025 5GR-DTAD: A Domain and Data-Driven Framework for Diagnosing Abnormal Downlink Throughput in 5G RAN
abstract
The advent of 5G wireless technology marks a significant milestone in telecommunications, enhancing consumer and industrial applications through Industry 4.0 technologies. The radio access network (RAN) and its downlink throughput are vital for Internet service providers. However, ensuring 5G RAN downlink throughput reliability requires robust failure diagnosis strategies for quick identification and resolution, posing challenges in cost efficiency, expert trustworthiness, and adaptability. We present 5G RAN Downlink Throughput Abnormality Diagnosis (5GR-DTAD), a diagnostic framework that fusing domain expertise with machine learning techniques, incorporating both feed-forward and long short-term memory (LSTM) neural networks. Experimental results on real-world data from multiple base stations show that 5GR-DTAD outperforms existing methods in precision, recall, and F1 scores by up to 29.85%, 42.28%, and 35.74%, respectively. 5GR-DTAD improves diagnostic accuracy while minimizing reliance on labeled data, offering an adaptable and cost-effective solution for various 5G RAN conditions.
Yuqian Yang, Cong Zhao 0001, Shusen Yang, Zongben Xu
IEEE Trans. Ind. Informatics4
2025 Anatomy-Aware Deep Unrolling for Task-Oriented Acceleration of Multi-Contrast MRI
abstract
Multi-contrast magnetic resonance imaging (MC-MRI) plays a crucial role in clinical practice. However, its performance is hindered by long scanning times and the isolation between image acquisition and downstream clinical diagnoses/treatments. Despite the activated research on accelerated MC-MRI, few existing studies prioritize personalized imaging tailored to individual patient characteristics and clinical needs. That is, the current approach often aims to enhance overall image quality, disregarding the specific pathologies or anatomical regions that are of particular interest to clinicians. To tackle this challenge, we propose an anatomy-aware unrolling-based deep network, dubbed as $\text {A}^{{2}}$ MC-MRI, offering promising interpretability and learning capacity for fast MC-MRI catering to downstream clinical needs. The network is unfolded from the iterative algorithm designed for a task-oriented MC-MRI reconstruction model. Specifically, to enhance concurrent MC-MRI of specific targets of interest (TOIs), the model integrates a learnable group sparsity with an anatomy-aware denoising prior. Within the anatomy-aware denoising prior, a segmentation network is involved to provide critical location information for TOI-enhanced denoising. Finally, such an unrolled network is jointly learned with k-space sampling patterns for task-oriented MC-MR reconstruction. Comprehensive evaluations on two public benchmarks as well as an in-house dataset demonstrate that our ${A}^{{2}}$ MC-MRI led to state-of-the-art performance in MC-MRI reconstruction under high acceleration rates, featuring notable enhancements in TOI imaging quality. The code will be available at https://github.com/ladderlab-xjtu/A2MC-MRI.
Yuzhu He, Chunfeng Lian, Ruyi Xiao, Fangmao Ju, Chao Zou, Zongben Xu, Jianhua Ma 0001
IEEE Trans. Medical Imaging6
2025 Domain-Generalized Discrete Diffusion Model for Cross-Domain Medical Image Segmentation
abstract
Domain shift is a significant challenge in medical image segmentation, primarily due to variations in image acquisition protocols, modalities, etc. Domain shift often causes models trained on a source domain to perform poorly on unseen target domains. In this work, we introduce the Domain-Generalized Discrete Diffusion Model for Segmentation (DG-DDM-Seg), a diffusion-based generative model designed for single-source domain generalization in medical image segmentation. DG-DDM-Seg generates discrete conditional distributions of segmentation masks. To ensure domain independence, we employ two key strategies: 1) We extract robust features from conditional images to enhance the domain independence of diffusion model. 2) We use both conditional images and pseudo-labels as inputs to improve cross-domain segmentation performance. Along this idea, we propose a two-path reverse diffusion process during training, utilizing Robust Feature Extraction Subnet and Mask-Generation Transformer to learn a domain-generalized discrete conditional distribution based on robust image features and pseudo-labels. This learned distribution is then used to generate segmentation masks for unseen target domains. Experimental results demonstrate that DG-DDM-Seg achieves state-of-the-art performance in cross-domain medical image segmentation, with domain shifts in modality, sequence, and site. The code is available at https://github.com/HeranYang/DG-DDM-Seg.
Heran Yang, Wenbo Hua, Zongben Xu, Jian Sun 0009
IEEE Trans. Medical Imaging3
2025 Improve Noise Tolerance of Robust Loss via Noise-Awareness
abstract
Robust loss minimization is an important strategy for handling robust learning issue on noisy labels. Current approaches for designing robust losses involve the introduction of noise-robust factors, i.e., hyperparameters, to control the trade-off between noise robustness and learnability. However, finding suitable hyperparameters for different datasets with noisy labels is a challenging and time-consuming task. Moreover, existing robust loss methods usually assume that all training samples share common hyperparameters, which are independent of instances. This limits the ability of these methods to distinguish the individual noise properties of different samples and overlooks the varying contributions of diverse training samples in helping models understand underlying patterns. To address above issues, we propose to assemble robust loss with instance-dependent hyperparameters to improve their noise tolerance with theoretical guarantee. To achieve setting such instance-dependent hyperparameters for robust loss, we propose a meta-learning method which is capable of adaptively learning a hyperparameter prediction function, called noise-aware-robust-loss-adjuster (NARL-Adjuster). Through mutual amelioration between hyperparameter prediction function and classifier parameters in our method, both of them can be simultaneously finely ameliorated and coordinated to attain solutions with good generalization capability. Four SOTA robust loss functions are attempted to be integrated with our algorithm, and comprehensive experiments substantiate the general availability and effectiveness of the proposed method in both its noise tolerance and performance. Meanwhile, the explicit parameterized structure makes the meta-learned prediction function ready to be transferrable and plug-and-play to unseen datasets with noisy labels. Specifically, we transfer our meta-learned NARL-Adjuster to unseen tasks, including several real noisy datasets, and achieve better performance compared with conventional hyperparameter tuning strategy, even with carefully tuned hyperparameters.
Kehui Ding, Deyu Meng, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2025 TRG-Net: An Interpretable and Controllable Rain Generator
abstract
Exploring and modeling the rain generation mechanism is critical for augmenting paired data to ease the training of rainy image processing models. Most of the conventional methods handle this task in an artificial physical rendering manner, through elaborately designing fundamental elements constituting rains. These kinds of methods, however, are over-dependent on human subjectivity, which limits their adaptability to real rains. In contrast, recent deep learning (DL) methods have achieved great success by training a neural network-based generator from pre-collected rainy image data. However, current methods usually design the generator in a "closed box" manner, increasing the learning difficulty and data requirements. To address these issues, this study proposes a novel DL-based rain generator, which fully takes the physical generation mechanism underlying rains into consideration and well encodes the learning of the fundamental rain factors (i.e., shape, orientation, length, width, and sparsity) explicitly into the deep network. Its significance lies in that the generator not only elaborately designs essential elements of the rain to simulate expected rains, like conventional artificial strategies, but also finely adapts to complicated and diverse practical rainy images, like DL methods. By rationally adopting the filter parameterization technique, the proposed rain generator is finely controllable with respect to rain factors and able to learn the distribution of these factors purely from data without the need for rain factor labels. Our unpaired generation experiments demonstrate that the rain generated by the proposed rain generator is not only of higher quality but also more effective for deraining and downstream tasks compared to current state-of-the-art rain generation methods. Besides, the paired data augmentation experiments, including both in-distribution and out-of-distribution (OOD), further validate the diversity of samples generated by our model for in-distribution deraining and OOD generalization tasks.
Zhiqiang Pang, Hong Wang 0021, Qi Xie 0002, Deyu Meng, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.5
2025 Efficient MU-MIMO Beamforming Based on Majorization-Minimization and Deep Unfolding
abstract
To release the full potentials of massive multi-user multiple-input multiple-output (MU-MIMO) for wireless communication, beamforming design is a must. In this paper, three algorithms are progressively developed for the maximization of the weighted sum-rate (WSR) problem. First, an effective beamforming algorithm is developed by applying the majorization-minimization (MM) procedure in two stages, by which the WSR problem is transformed into a series of quadratic convex problems. We prove that the proposed algorithm converges to a stationary point of the WSR. Second, to improve the efficiency of the two-stage beamforming algorithm, the inherent low-dimensional structure within the beamforming update is exploited aiming to reduce the computational complexity of the matrix inversion. Third, to further reduce the complexity, a deep unfolding beamforming network is developed, which unfolds the improved beamforming algorithm into a layer-wise structure and employs a trainable module structured on the dynamics developed to approximate the matrix inversion. Experimental results demonstrate that the proposed algorithms perform significantly better than the classical weighted minimum mean square error (WMMSE) beamforming and state-of-the-art deep unfolding beamformers in terms of sum-rate and require significantly less CPU time.
Qian Xu 0017, Jianyong Sun, Zongben Xu
IEEE Trans. Wirel. Commun.3
2025 Deep Learning-Empowered Secure Predictive Beamforming Design for Integrated Sensing and Communications Systems
abstract
In the era of upcoming sixth-generation (6G) wireless systems, the intelligent integrated sensing and communication (ISAC) paradigm has emerged as a pivotal research domain, catalyzing advancement across a wide range of applications. In this paper, we investigate an ISAC-assisted anti-eavesdropping communication system, where an ISAC ground base station exploits its radar function to track potential aerial eavesdroppers and implements predictive beamforming to ensure secure communications with multiple ground users. We harness the powerful capability of the Transformer for time series prediction to establish a novel deep neural network, termed the ISACformer, for constructing predictive beamformers via exploiting previously estimated channel state information in an unsupervised manner. By eliminating the need for explicit channel prediction, our proposed framework effectively reduces signaling overhead and complexity. In addition, by formulating a weighted objective function, our design meticulously balances the trade-off between the ergodic achievable worst-case secrecy rate for ground users and the ergodic Cramér-Rao lower bound for the kinematic parameters of potential aerial eavesdroppers. Simulation results demonstrate that the proposed ISACformer can deliver the desired predictive beamforming for harmonizing radar and communication functionalities effectively. Moreover, our method achieves performance approaching the theoretical upper bound obtained by ignoring multi-user interference, thereby highlighting the robustness of the proposed approach.
Zhen Qiao, Faheem Ahmad Khan, Guanzhang Liu, Zhiqiang Wei 0001, Jiang Xue 0001, Zongben Xu, Derrick Wing Kwan Ng
IEEE Trans. Wirel. Commun.7
2025 Cross-Channel Model-Driven Learning for Massive MIMO Detection by HyperNetwork
abstract
For the signal detection problem in a multiple-input multiple-output (MIMO) system, it has been demonstrated that deep learning can improve the detection accuracy and/or reduce the complexity of traditional detection algorithms under the assumption that the channel scenario remains the same in training and test. However, this assumption is not appropriate since the communication environment in practice is constantly changing. As a result, the performance of deep-learning-based detection methods will degrade significantly due to their lack of generalization ability. To address this problem, we model the channel scenario adaptation problem as a multi-scenario learning task and propose two schemes to improve the adaptability of model-driven detection network to cross-channel scenarios. For the case where the test channel scenario has been seen in the training stage, a hypernetwork is introduced to the deep-learning-based iterative soft thresholding algorithm (DISTA) to generate a personalized set of network parameters for each channel scenario, which is named hyperDISTA. Experimental results show that hyperDISTA trained in multiple channel scenarios can not only adapt to each seen channel scenario but also outperform existing deep-learning-based detectors trained in the single channel scenario at high signal-to-noise ratio (SNR) regimes. For the case where the test channel scenario is unseen in the training stage, we propose to retrain the hyperDISTA in a semi-supervised manner. Experimental results show that the retrained hyperDISTA achieves a performance that is comparable to that of the maximum likelihood detection algorithm (MLD).
Yiqing Zhang 0001, Jianyong Sun, Jiang Xue 0001, Zongben Xu
IEEE Trans. Wirel. Commun.4
2024 Blind Proximal Diffusion Model for Joint Image and Sensitivity Estimation in Parallel MRI
Xing Li 0027, Yan Yang 0007, Hairong Zheng, Zongben Xu
MICCAI (7)4
2024 Adversarial data splitting for domain generalization
Xiang Gu 0005, Jian Sun 0009, Zongben Xu
Sci. China Inf. Sci.3
2024 Adversarial Reweighting with α-Power Maximization for Domain Adaptation
Xiang Gu 0005, Yan Yang 0007, Jian Sun 0009, Zongben Xu
Int. J. Comput. Vis.5
2024 Joint orthogonal symmetric non-negative matrix factorization for community detection in attribute network
Qingming Kong, Jianyong Sun, Zongben Xu
Knowl. Based Syst.3
2024 Rotation Equivariant Proximal Operator for Deep Unfolding Methods in Image Restoration
abstract
The deep unfolding approach has attracted significant attention in computer vision tasks, which well connects conventional image processing modeling manners with more recent deep learning techniques. Specifically, by establishing a direct correspondence between algorithm operators at each implementation step and network modules within each layer, one can rationally construct an almost "white box" network architecture with high interpretability. In this architecture, only the predefined component of the proximal operator, known as a proximal network, needs manual configuration, enabling the network to automatically extract intrinsic image priors in a data-driven manner. In current deep unfolding methods, such a proximal network is generally designed as a CNN architecture, whose necessity has been proven by a recent theory. That is, CNN structure substantially delivers the translational symmetry image prior, which is the most universally possessed structural prior across various types of images. However, standard CNN-based proximal networks have essential limitations in capturing the rotation symmetry prior, another universal structural prior underlying general images. This leaves a large room for further performance improvement in deep unfolding approaches. To address this issue, this study makes efforts to suggest a high-accuracy rotation equivariant proximal network that effectively embeds rotation symmetry priors into the deep unfolding framework. Especially, we deduce, for the first time, the theoretical equivariant error for such a designed proximal network with arbitrary layers under arbitrary rotation degrees. This analysis should be the most refined theoretical conclusion for such error evaluation to date and is also indispensable for supporting the rationale behind such networks with intrinsic interpretability requirements. Through experimental validation on different vision tasks, including blind image super-resolution, medical image reconstruction, and image de-raining, the proposed method is validated to be capable of directly replacing the proximal network in current deep unfolding architecture and readily enhancing their state-of-the-art performance. This indicates its potential usability in general vision tasks.
Jiahong Fu, Qi Xie 0002, Deyu Meng, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Unsupervised and Semi-Supervised Robust Spherical Space Domain Adaptation
abstract
Adversarial domain adaptation has been an effective approach for learning domain-invariant features by adversarial training. In this paper, we propose a novel adversarial domain adaptation approach defined in the spherical feature space, in which we define spherical classifier for label prediction and spherical domain discriminator for discriminating domain labels. In the spherical feature space, we develop a spherical robust pseudo-label loss to utilize pseudo-labels robustly, which weights the importance of the estimated labels of target domain data by the posterior probability of correct labeling, modeled by the Gaussian-uniform mixture model in the spherical space. Our proposed approach can be generally applied to both unsupervised and semi-supervised domain adaptation settings. In particular, to tackle the semi-supervised domain adaptation setting where a few labeled target domain data are available for training, we propose a novel reweighted adversarial training strategy for effectively reducing the intra-domain discrepancy within the target domain. We also present theoretical analysis for the proposed method based on the domain adaptation theory. Extensive experiments are conducted on multiple benchmarks for object recognition, digit recognition, and face recognition. The results show that our method either surpasses or is competitive compared with the recent methods for both unsupervised and semi-supervised domain adaptation. Ablation studies also confirm the effectiveness of the spherical classifier, spherical discriminator, spherical robust pseudo-label loss, and reweighted adversarial training strategy.
Xiang Gu 0005, Jian Sun 0009, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 ISP-IRLNet: Joint optimization of interpretable sampler and implicit regularization learning network for accerlerated MRI
Xing Li 0027, Yan Yang 0007, Hairong Zheng, Zongben Xu
Pattern Recognit.4
2024 Probabilistic Matrix Factorization for Data With Attributes Based on Finite Mixture Modeling
abstract
Matrix factorization (MF) methods decompose a data matrix into a product of two-factor matrices (denoted as U and V ) which are with low ranks. In this article, we propose a generative latent variable model for the data matrix, in which each entry is assumed to be a Gaussian with mean to be the inner product of the corresponding columns of U and V . The prior of each column of U and V is assumed to be as a finite mixture of Gaussians. Further, we propose to model the attribute matrix with the data matrix jointly by considering them as conditional independence with respect to the factor matrix U , building upon previously defined model for the data matrix. Due to the intractability of the proposed models, we employ variational Bayes to infer the posteriors of the factor matrices and the clustering relationships, and to optimize for the model parameters. In our development, the posteriors and model parameters can be readily computed in closed forms, which is much more computationally efficient than existing sampling-based probabilistic MF models. Comprehensive experimental studies of the proposed methods on collaborative filtering and community detection tasks demonstrate that the proposed methods achieve the state-of-the-art performance against a great number of MF-based and non-MF-based algorithms.
Qingming Kong, Jianyong Sun, Zongben Xu
IEEE Trans. Cybern.4
2024 Deep Learning Accelerated Blind Seismic Acoustic-Impedance Inversion
abstract
Blind seismic acoustic-impedance (AI) inversion is a technique for obtaining the AI of the subsurface medium without given a wavelet. An effective way to solve the blind inversion problem is to split the multi-parameter problem into two single-parameter subproblems and solve them in an alternative iteration way. However, this method becomes time-consuming when dealing with large-scale 3D problems and faces challenges in selecting suitable regularization parameters. To overcome these shortcomings, we propose a deep learning accelerated blind seismic AI inversion (DLA-BSAII) method. It mainly has three steps: (1) Only a few 2D profiles are selected from the whole 3D data, and their corresponding AI models and wavelets are inverted using conventional blind seismic AI inversion method. (2) The results of the first step are used to train deep networks to realize the nonlinear mapping from a 2D seismic profile to AI and wavelet. In addition, the trained deep networks are used to generate predictions of AI models and wavelets for the remaining 2D profiles. (3) Benefiting from the predicted AI models and wavelets, a new alternative iteration method with fewer but more effective regularization terms is proposed to obtain the final inverted AI models and wavelets of the remaining 2D profiles. It has the advantages of easier selection of regularization parameters and faster convergence speed. Synthetic and field data examples verify that DLA-BSAII outperforms conventional methods in terms of both efficiency and inversion accuracy.
Zhaoqi Gao, Meiqian Guo, Chuang Li 0003, Zhen Li 0016, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.6
2024 Robust Saliency-Driven Quality Adaptation for Mobile 360-Degree Video Streaming
abstract
Mobile 360-degree video streaming has grown significantly in popularity but the quality of experience (QoE) suffers from insufficient and variable wireless network bandwidth. Recently, saliency-driven 360-degree streaming overcomes the buffer size limitation of head movement trajectory (HMT)-driven solutions and thus strikes a better balance between video quality and rebuffering. However, inaccurate network estimations and intrinsic saliency bias still challenge saliency-based streaming approaches, limiting further QoE improvement. To address these challenges, we design a robust saliency-driven quality adaptation algorithm for 360-degree video streaming, RoSal360. Specifically, we present a practical, tile-size-aware deep neural network (DNN) model with a decoupled self-attention architecture to accurately and efficiently predict the transmission time of video tiles. Moreover, we design a reinforcement learning (RL)-driven online correction algorithm to robustly compensate the improper quality allocations due to saliency bias. Through extensive prototype evaluations over real wireless network environments including commodity WiFi, 4G/LTE, and 5G links in the wild, RoSal360 significantly enhances the video quality and reduces the rebuffering ratio, thereby improving the viewer QoE, compared to the state-of-the-art algorithms.
Shibo Wang 0002, Shusen Yang, Hairong Su, Cong Zhao 0001, Chenren Xu, Feng Qian 0001, Nanbin Wang, Zongben Xu
IEEE Trans. Mob. Comput.8
2024 Noise-Generating and Imaging Mechanism Inspired Implicit Regularization Learning Network for Low Dose CT Reconstrution
abstract
Low-dose computed tomography (LDCT) helps to reduce radiation risks in CT scanning while maintaining image quality, which involves a consistent pursuit of lower incident rays and higher reconstruction performance. Although deep learning approaches have achieved encouraging success in LDCT reconstruction, most of them treat the task as a general inverse problem in either the image domain or the dual (sinogram and image) domains. Such frameworks have not considered the original noise generation of the projection data and suffer from limited performance improvement for the LDCT task. In this paper, we propose a novel reconstruction model based on noise-generating and imaging mechanism in full-domain, which fully considers the statistical properties of intrinsic noises in LDCT and prior information in sinogram and image domains. To solve the model, we propose an optimization algorithm based on the proximal gradient technique. Specifically, we derive the approximate solutions of the integer programming problem on the projection data theoretically. Instead of hand-crafting the sinogram and image regularizers, we propose to unroll the optimization algorithm to be a deep network. The network implicitly learns the proximal operators of sinogram and image regularizers with two deep neural networks, providing a more interpretable and effective reconstruction procedure. Numerical results demonstrate our proposed method improvements of > 2.9 dB in peak signal to noise ratio, > 1.4% promotion in structural similarity metric, and > 9 HU decrements in root mean square error over current state-of-the-art LDCT methods.
Xing Li 0027, Kaili Jing, Yan Yang 0007, Jianhua Ma 0001, Hairong Zheng, Zongben Xu
IEEE Trans. Medical Imaging7
2024 Direction-of-Arrival Estimation for Constant Modulus Signals Using a Structured Matrix Recovery Technique
abstract
This paper addresses the problem of direction-of-arrival (DOA) estimation for constant modulus (CM) source signals using a uniform or sparse linear array. Existing methods typically exploit either the Vandermonde structure of the steering matrix or the CM structure of source signals only. In this paper, we propose a structuredmatrix recovery technique (SMART) for CM DOA estimation via fully exploiting the two structures. In particular, we reformulate the highly nonconvex CM DOA estimation problems in the noiseless and noisy cases as equivalent rank-constrained Hankel-Toeplitz matrix recovery problems, in which the Vandermonde structure is captured by a series of Hankel-Toeplitz block matrices, of which the number equals the number of snapshots, and the CM structure is guaranteed by letting the block matrices share a same Toeplitz submatrix. The alternating direction method of multipliers (ADMM) is applied to solve the resulting rank-constrained problems and the DOAs are uniquely retrieved from the numerical solution. Extensive simulations are carried out to corroborate our analysis and confirm that the proposed SMART outperforms state-of-the-art algorithms in terms of the maximum number of locatable sources and statistical efficiency.
Xunmeng Wu, Zai Yang, Zhiqiang Wei 0001, Zongben Xu
IEEE Trans. Wirel. Commun.4
2023 Spectral Super-Resolution on the Unit Circle Via Gradient Descent
abstract
We study the spectral super-resolution problem, which concerns the construction of an undamped spectrally sparse signal and its frequencies from its partially revealed entries. We propose a nonconvex method composed of a Hankel-Toeplitz matrix factorization model and a gradient descent algorithm termed as HT-GD. The model is equivalent to an ℓ0norm con-strained problem, which ensures that the all signal structures including the spectral poles lying on the unit circle are exploited. The gradient descent algorithm, consisting of spectral initialization and iterative refinement, is computationally efficient. Numerical results demonstrate that our method out-performs state-of-the-art approaches in terms of accuracy and computational speed.
Xunmeng Wu, Zai Yang, Jian-Feng Cai 0001, Zongben Xu
ICASSP4
2023 Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images
abstract
Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D face generation remains an open problem. In this paper, we propose a text-guided 3D faces generation method, refer as TG-3DFace, for generating realistic 3D faces using text guidance. Specifically, we adopt an unconditional 3D face generation framework and equip it with text conditions, which learns the text-guided 3D face generation with only text-2D face data. On top of that, we propose two text-to-face cross-modal alignment techniques, including the global contrastive learning and the fine-grained alignment module, to facilitate high semantic consistency between generated 3D faces and input texts. Besides, we present directional classifier guidance during the inference process, which encourages creativity for out-of-domain generations. Compared to the existing methods, TG-3DFace creates more realistic and aesthetically pleasing 3D faces, boosting 9% multi-view consistency (MVIC) over Latent3D. The rendered face images generated by TG-3DFace achieve higher FID and CLIP score than text-to-2D face/image generation models, demonstrating our superiority in generating realistic and semantic-consistent textures.
Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 0009, Xiaodan Liang, Huibin Li 0001, Zongben Xu, Songcen Xu, Wei Zhang 0196, Hang Xu 0004
ICCV7
2023 Optimal Transport-Guided Conditional Score-Based Diffusion Model
abstract
Conditional score-based diffusion model (SBDM) is for conditional generation of target data with paired data as condition, and has achieved great success in image translation. However, it requires the paired data as condition, and there would be insufficient paired data provided in real-world applications. To tackle the applications with partially paired or even unpaired dataset, we propose a novel Optimal Transport-guided Conditional Score-based diffusion model (OTCS) in this paper. We build the coupling relationship for the unpaired or partially paired dataset based on $L_2$-regularized unsupervised or semi-supervised optimal transport, respectively. Based on the coupling relationship, we develop the objective for training the conditional score-based model for unpaired or partially paired settings, which is based on a reformulation and generalization of the conditional SBDM for paired setting. With the estimated coupling relationship, we effectively train the conditional score-based model by designing a ``resampling-by-compatibility'' strategy to choose the sampled data with high compatibility as guidance. Extensive experiments on unpaired super-resolution and semi-paired image-to-image translation demonstrated the effectiveness of the proposed OTCS model. From the viewpoint of optimal transport, OTCS provides an approach to transport data across distributions, which is a challenge for OT on large-scale datasets. We theoretically prove that OTCS realizes the data transport in OT with a theoretical bound.
Xiang Gu 0005, Jian Sun 0009, Zongben Xu
NeurIPS4
2023 Robust channel estimation based on the maximum entropy principle
Zhengyang Hu 0001, Jiang Xue 0001, Feng Li 0057, Qian Zhao 0002, Deyu Meng, Zongben Xu
Sci. China Inf. Sci.6
2023 Learning unified mutation operator for differential evolution by natural evolution strategies
Haotian Zhang 0023, Jianyong Sun, Zongben Xu, Jialong Shi
Inf. Sci.3
2023 Deep expectation-maximization network for unsupervised image segmentation and clustering
Yannan Pu, Jian Sun 0009, Niansheng Tang, Zongben Xu
Image Vis. Comput.4
2023 Learning an Explicit Hyper-parameter Prediction Function Conditioned on Tasks
abstract
Meta learning has attracted much attention recently in machine learning community. Contrary to conventional machine learning aiming to learn inherent prediction rules to predict labels for new query data, meta learning aims to learn the learning methodology for machine learning from observed tasks, so as to generalize to new query tasks by leveraging the meta-learned learning methodology. In this study, we achieve such learning methodology by learning an explicit hyper-parameter prediction function shared by all training tasks, and we call this learning process as Simulating Learning Methodology (SLeM). Specifically, this function is represented as a parameterized function called meta-learner, mapping from a training/test task to its suitable hyper-parameter setting, extracted from a pre-specified function set called meta learning machine. Such setting guarantees that the meta-learned learning methodology is able to flexibly fit diverse query tasks, instead of only obtaining fixed hyper-parameters by many current meta learning methods, with less adaptability to query task's variations. Such understanding of meta learning also makes it easily succeed from traditional learning theory for analyzing its generalization bounds with general losses/tasks/models. The theory naturally leads to some feasible controlling strategies for ameliorating the quality of the extracted meta-learner, verified to be able to finely ameliorate its generalization capability in some typical meta learning applications, including few-shot regression, few-shot classification and domain generalization. The source code of our method is released at https://github.com/xjtushujun/SLeM-Theory.
Deyu Meng, Zongben Xu
J. Mach. Learn. Res.3
2023 Variational Data-Free Knowledge Distillation for Continual Learning
abstract
Deep neural networks suffer from catastrophic forgetting when trained on sequential tasks in continual learning. Various methods rely on storing data of previous tasks to mitigate catastrophic forgetting, which is prohibited in real-world applications considering privacy and security issues. In this paper, we consider a realistic setting of continual learning, where training data of previous tasks are unavailable and memory resources are limited. We contribute a novel knowledge distillation-based method in an information-theoretic framework by maximizing mutual information between outputs of previously learned and current networks. Due to the intractability of computation of mutual information, we instead maximize its variational lower bound, where the covariance of variational distribution is modeled by a graph convolutional network. The inaccessibility of data of previous tasks is tackled by Taylor expansion, yielding a novel regularizer in network training loss for continual learning. The regularizer relies on compressed gradients of network parameters. It avoids storing previous task data and previously learned networks. Additionally, we employ self-supervised learning technique for learning effective features, which improves the performance of continual learning. We conduct extensive experiments including image classification and semantic segmentation, and the results show that our method achieves state-of-the-art performance on continual learning benchmarks.
Xiaorong Li, Shipeng Wang 0002, Jian Sun 0009, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 CMW-Net: Learning a Class-Aware Sample Weighting Mapping for Robust Deep Learning
abstract
Modern deep neural networks can easily overfit to biased training data containing corrupted labels or class imbalance. Sample re-weighting methods are popularly used to alleviate this data bias issue. Most current methods, however, require to manually pre-specify the weighting schemes relying on the characteristics of the investigated problem and training data. This makes them fairly hard to be generally applied in practical scenarios, due to their significant complexities and inter-class variations of data bias. To address this issue, we propose a meta-model capable of adaptively learning an explicit weighting scheme directly from data. Specifically, by seeing each training class as a separate learning task, our method aims to extract an explicit weighting function with sample loss and task/class feature as input, and sample weight as output, expecting to impose adaptively varying weighting schemes to different sample classes based on their own intrinsic bias characteristics. Extensive experiments substantiate the capability of our method on achieving proper weighting schemes in various data bias cases, like class imbalance, feature-independent and dependent label noises, and more complicated bias scenarios beyond conventional cases. Besides, the task-transferability of the learned weighting scheme is also substantiated, by readily deploying the weighting function learned on relatively smaller-scale CIFAR-10 dataset on much larger-scale full WebVision dataset. The general availability of our method for multiple robust deep learning issues, including partial-label learning, semi-supervised learning and selective classification, has also been validated. Code for reproducing our experiments is available at https://github.com/xjtushujun/CMW-Net.
Deyu Meng, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 MLR-SNet: Transferable LR Schedules for Heterogeneous Tasks
abstract
The learning rate (LR) is one of the most important hyperparameters in stochastic gradient descent (SGD) algorithm for training deep neural networks (DNN). However, current hand-designed LR schedules need to manually pre-specify a fixed form, which limits their ability to adapt to practical non-convex optimization problems due to the significant diversification of training dynamics. Meanwhile, it always needs to search proper LR schedules from scratch for new tasks, which, however, are often largely different with task variations, like data modalities, network architectures, or training data capacities. To address this learning-rate-schedule setting issue, we propose to parameterize LR schedules with an explicit mapping formulation, called MLR-SNet. The learnable parameterized structure brings more flexibility for MLR-SNet to learn a proper LR schedule to comply with the training dynamics of DNN. Image and text classification benchmark experiments substantiate the capability of our method for achieving proper LR schedules. Moreover, the explicit parameterized structure makes the meta-learned LR schedules capable of being transferable and plug-and-play, which can be easily generalized to new heterogeneous tasks. We transfer our meta-learned MLR-SNet to query tasks like different training epochs, network architectures, data modalities, dataset sizes from the training ones, and achieve comparable or even better performance compared with hand-designed LR schedules specifically designed for the query tasks. The robustness of MLR-SNet is also substantiated when the training data are biased with corrupted noise. We further prove the convergence of the SGD algorithm equipped with LR schedule produced by our MLR-SNet, with the convergence rate comparable to the best-known ones of the algorithm for solving the problem. The source code of our method is released at https://github.com/xjtushujun/MLR-SNet.
Yanwen Zhu, Qian Zhao 0002, Deyu Meng, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Fourier Series Expansion Based Filter Parametrization for Equivariant Convolutions
abstract
It has been shown that equivariant convolution is very helpful for many types of computer vision tasks. Recently, the 2D filter parametrization technique has played an important role for designing equivariant convolutions, and has achieved success in making use of rotation symmetry of images. However, the current filter parametrization strategy still has its evident drawbacks, where the most critical one lies in the accuracy problem of filter representation. To address this issue, in this paper we explore an ameliorated Fourier series expansion for 2D filters, and propose a new filter parametrization method based on it. The proposed filter parametrization method not only finely represents 2D filters with zero error when the filter is not rotated (similar as the classical Fourier series expansion), but also substantially alleviates the aliasing-effect-caused quality degradation when the filter is rotated (which usually arises in classical Fourier series expansion method). Accordingly, we construct a new equivariant convolution method based on the proposed filter parametrization method, named F-Conv. We prove that the equivariance of the proposed F-Conv is exact in the continuous domain, which becomes approximate only after discretization. Moreover, we provide theoretical error analysis for the case when the equivariance is approximate, showing that the approximation error is related to the mesh size and filter size. Extensive experiments show the superiority of the proposed method. Particularly, we adopt rotation equivariant convolution methods to a typical low-level image processing task, image super-resolution. It can be substantiated that the proposed F-Conv based method evidently outperforms classical convolution based methods. Compared with pervious filter parametrization based methods, the F-Conv performs more accurately on this low-level image processing task, reflecting its intrinsic capability of faithfully preserving rotation symmetries in local image features.
Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Memory efficient data-free distillation for continual learning
Xiaorong Li, Shipeng Wang 0002, Jian Sun 0009, Zongben Xu
Pattern Recognit.4
2023 Meta-learning-based adversarial training for deep 3D face recognition on point clouds
Cuican Yu, Huibin Li 0001, Jian Sun 0009, Zongben Xu
Pattern Recognit.5
2023 An Unrolled Implicit Regularization Network for Joint Image and Sensitivity Estimation in Parallel MR Imaging with Convergence Guarantee
abstract
Abstract. Parallel imaging (PI), relying on multicoils to sense [Formula: see text]-space data, is an effective technique to accelerate magnetic resonance imaging by exploiting spatial sensitivity coding of multiple coils, with an integrated compressive sensing (CS) technology to achieve higher acceleration. In this paper, we propose a novel nonconvex reconstruction model and its proximal alternating linearized minimization (PALM) algorithm for PI in a blind setting that MR image and multichannel sensitivity maps are jointly estimated, regularized by image and sensitivity regularizers. Instead of hand-crafting the image and sensitivity regularizers, we propose unrolling the PALM algorithm to be a deep network for Blind Parallel MRI, dubbed as BPMRI-Net, with two learnable subnetworks to substitute the proximal operators of the image and sensitivity regularizers. We theoretically prove the linear convergence of BPMRI-Net as an iterative algorithm, which alternately updates two variables based on the learnable proximal operators. The learned BPMRI-Net can simultaneously output the MR image and sensitivity maps from undersampled multichannel [Formula: see text]-space data even when the number of low-frequency sampling lines in the center of [Formula: see text]-space is small. Numerical results demonstrate the effectiveness of our method with state-of-the-art reconstruction accuracy.
Yan Yang 0007, Yizhou Wang 0006, Jiazhen Wang, Jian Sun 0009, Zongben Xu
SIAM J. Imaging Sci.5
2023 Continuous Encoding for Overlapping Community Detection in Attributed Network
abstract
Detecting overlapping communities of an attribute network is a ubiquitous yet very difficult task, which can be modeled as a discrete optimization problem. Besides the topological structure of the network, node attributes and node overlapping aggravate the difficulty of community detection significantly. In this article, we propose a novel continuous encoding method to convert the discrete-natured detection problem to a continuous one by associating each edge and node attribute in the network with a continuous variable. Based on the encoding, we propose to solve the converted continuous problem by a multiobjective evolutionary algorithm (MOEA) based on decomposition. To find the overlapping nodes, a heuristic based on double-decoding is proposed, which is only with linear complexity. Furthermore, a postprocess community merging method in consideration of node attributes is developed to enhance the homogeneity of nodes in the detected communities. Various synthetic and real-world networks are used to verify the effectiveness of the proposed approach. The experimental results show that the proposed approach performs significantly better than a variety of evolutionary and nonevolutionary methods on most of the benchmark networks.
Wei Zheng 0004, Jianyong Sun, Qingfu Zhang 0001, Zongben Xu
IEEE Trans. Cybern.4
2023 Generating Azimuth-Reflection Angle Gathers From Reverse Time Migration Using the High-Dimensional Local Phase Space Approximation of Seismic Wavefields
abstract
Amplitude-preserving angle gathers are ideal inputs for seismic prestack inversion. However, due to the limitation of computational efficiency, generating subsurface azimuth–reflection angle gathers from 3-D seismic imaging is still a very difficult task. In this article, we propose a new method to generate azimuth–reflection angle gathers from 3-D reverse time migration (RTM). The proposed method approximately reconstructs the source wavefield using high-dimensional wavelets and the excitation information. After using directional vectors to calculate the subsurface observation angles and applying the cross correlation imaging condition, we can generate azimuth–reflection angle gathers by angle binning. Without storing source wavefields or reconstructing source wavefields using boundary conditions, the proposed method has high computational efficiency. Numerical experiments on a synthetic model and a real marine seismic dataset demonstrate that compared with the excitation amplitude imaging condition, the proposed method can generate azimuth–reflection angle gathers with continuous complete events and high signal-to-noise ratio. The image quality and resolution of angle gathers are significantly improved. At the same time, the computational complexity does not increase much.
Feipeng Li, Jinghuai Gao, Zhaoqi Gao, Chuang Li 0003, Qingzhen Wang, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.7
2023 Learning Unified Hyper-Network for Multi-Modal MR Image Synthesis and Tumor Segmentation With Missing Modalities
abstract
Accurate segmentation of brain tumors is of critical importance in clinical assessment and treatment planning, which requires multiple MR modalities providing complementary information. However, due to practical limits, one or more modalities may be missing in real scenarios. To tackle this problem, existing methods need to train multiple networks or a unified but fixed network for various possible missing modality cases, which leads to high computational burdens or sub-optimal performance. In this paper, we propose a unified and adaptive multi-modal MR image synthesis method, and further apply it to tumor segmentation with missing modalities. Based on the decomposition of multi-modal MR images into common and modality-specific features, we design a shared hyper-encoder for embedding each available modality into the feature space, a graph-attention-based fusion block to aggregate the features of available modalities to the fused features, and a shared hyper-decoder for image reconstruction. We also propose an adversarial common feature constraint to enforce the fused features to be in a common space. As for missing modality segmentation, we first conduct the feature-level and image-level completion using our synthesis method and then segment the tumors based on the completed MR images together with the extracted common features. Moreover, we design a hypernet-based modulation module to adaptively utilize the real and synthetic modalities. Experimental results suggest that our method can not only synthesize reasonable multi-modal MR images, but also achieve state-of-the-art performance on brain tumor segmentation with missing modalities.
Heran Yang, Jian Sun 0009, Zongben Xu
IEEE Trans. Medical Imaging3
2023 Understanding Deep MIMO Detection
abstract
Incorporating deep learning (DL) into multiple-input multiple-output (MIMO) detection has been deemed as a promising technique for future wireless communications. However, most of the DL-based MIMO detection algorithms are lack of interpretation on internal mechanisms. In this paper, we analyze the performance of the DL-based MIMO detection to better understand its strengths and weaknesses. We investigate and compare two different models: data-driven DL detector with neural networks activated by rectifier linear unit (ReLU) function and model-driven DL detector based on traditional detection algorithms. We show that the data-driven DL detector asymptotically approaches to the maximum a posterior (MAP) detector in various scenarios but requires a large amount of training samples to converge in time-varying channels. On the other hand, the model-driven DL detector utilizes the expert knowledge to alleviate the impact of channels and achieves relatively high detection accuracy with a small set of training data. Simulation results confirm our analytical results and demonstrate the effectiveness of the DL-based MIMO detection for both linear and nonlinear signal systems.
Feifei Gao 0001, Hao Zhang 0026, Geoffrey Ye Li, Zongben Xu
IEEE Trans. Wirel. Commun.5
2023 MIMO Detector Selection With Federated Learning
abstract
In this paper, we develop a dynamic detection network (DDNet) based detector for multiple-input multiple-output (MIMO) systems. By constructing an improved DetNet (IDetNet) detector and the OAMPNet detector as two independent network branches, the DDNet detector performs sample-wise dynamic routing to adaptively select a better one between the IDetNet and the OAMPNet detectors for every samples under different system conditions. To avoid the prohibitive transmission overhead of dataset collection in centralized learning (CL), we propose the federated averaging (FedAve)-DDNet detector, where all raw data are kept at local clients and only locally trained model parameters are transmitted to the central server for aggregation. To further reduce the transmission overhead, we develop the federated gradient sparsification (FedGS)-DDNet detector by randomly sampling gradients with elaborately calculated probability when uploading gradients to the central server. Based on simulation results, the proposed DDNet detector consistently outperforms other detectors under all system conditions thanks to the sample-wise dynamic routing. Moreover, the federated DDNet detectors, especially the FedGS-DDNet detector, can reduce the transmission overhead by at least 25.7% while maintaining satisfactory detection accuracy.
Yuwen Yang, Feifei Gao 0001, Jiang Xue 0001, Zongben Xu
IEEE Trans. Wirel. Commun.5
2022 KXNet: A Model-Driven Deep Neural Network for Blind Super-Resolution
Jiahong Fu, Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu
ECCV (19)6
2022 SalientVR: saliency-driven mobile 360-degree video streaming with gaze information
abstract
Mobile 360° video streaming has grown significantly in popularity but the quality of experience (QoE) suffers from insufficient wireless network bandwidth. The state-of-the-art solutions are limited by the temporal correlation assumption. Recent studies are aware of the potential of saliency to further QoE improvement, but several fundamental challenges about saliency judgment, saliency acquirement, and quality adaptation are still not fully addressed. To solve these challenges, we present SalientVR, a saliency-driven mobile 360° video streaming system integrated with gaze information. We design (i) a precise gaze-driven saliency judging criterion for mobile VR viewers, (ii) two pragmatic gaze-driven, tile-level saliency acquiring methods based on cross-user similarity and a specific content-aware deep neural network respectively, and (iii) a lightweight saliency-aware quality adaptation algorithm with a motion-assisted online correction, which is robust to wireless bandwidth vagaries and saliency bias. Moreover, we contribute a gaze-annotated dataset and a gaze-driven quality assessment metric for 360° videos. By extensive prototype evaluations (based on dataset tests and user studies), compared to alternatives, SalientVR significantly enhances the video quality and reduces the rebuffering ratio over 4G/LTE network emulations and in the wild, which achieves a 43.68% QoE improvement.
Shibo Wang 0002, Shusen Yang, Chenren Xu, Feng Qian 0001, Nanbin Wang, Zongben Xu
MobiCom9
2022 Keypoint-Guided Optimal Transport with Applications in Heterogeneous Domain Adaptation
abstract
Existing Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimization, which may cause incorrect matching in some cases. In many applications, annotating a few matched keypoints across domains is reasonable or even effortless in annotation burden. It is valuable to investigate how to leverage the annotated keypoints to guide the correct matching in OT. In this paper, we propose a novel KeyPoint-Guided model by ReLation preservation (KPG-RL) that searches for the matching guided by the keypoints in OT. To impose the keypoints in OT, first, we propose a mask-based constraint of the transport plan that preserves the matching of keypoint pairs. Second, we propose to preserve the relation of each data point to the keypoints to guide the matching. The proposed KPG-RL model can be solved by the Sinkhorn's algorithm and is applicable even when distributions are supported in different spaces. We further utilize the relation preservation constraint in the Kantorovich Problem and Gromov-Wasserstein model to impose the guidance of keypoints in them. Meanwhile, the proposed KPG-RL model is extended to partial OT setting. As an application, we apply the proposed KPG-RL model to the heterogeneous domain adaptation. Experiments verified the effectiveness of the KPG-RL model.
Xiang Gu 0005, Yucheng Yang 0004, Jian Sun 0009, Zongben Xu
NeurIPS5
2022 LDP-IDS: Local Differential Privacy for Infinite Data Streams
abstract
Local differential privacy (LDP) is promising for private streaming data collection and analysis. However, existing few LDP studies over streams either apply to finite streams only or may suffer from insufficient protection. This paper investigates this problem by proposing LDP-IDS, a novel w-event LDP paradigm to provide practical privacy guarantee for infinite streams. By constructing a unified error analysis, we adapt the existing budget division framework in centralized differential privacy (CDP) for LDP-IDS, which however incurs prohibitive noise and expensive communication cost. To this end, we propose a novel and extensible framework of population division and recycling, as well as online adaptive population division algorithms for LDP-IDS. We provide theoretical guarantees and demonstrate, through extensive discussions, that our proposed framework not only achieves significant reduction in utility loss and communication overhead, but also enjoys great compatibility for varied analytic tasks and flexibility of incorporating ideas of many existing stream algorithms. Extensive experiments on synthetic and real-world datasets validate the high effectiveness, efficiency, and flexibility of our proposed framework and methods.
Xuebin Ren, Weiren Yu, Shusen Yang, Cong Zhao 0001, Zongben Xu
SIGMOD Conference6
2022 Unified Mathematical Framework for Intelligent Transceiver Design
abstract
This paper proposes a unified mathematical frame-work for intelligent transceiver design. It mainly includes three most important modules in the communication system, namely, beamforming, channel estimation and Multiple-Input Multiple-Output (MIMO) detection. Firstly, the mathematical correlation behind different algorithms of a single communication module is analyzed, the purpose is to realize the unification between different algorithms of a specific communication module. Next, a cross-module unified mathematical framework is proposed. Finally, an AI architecture for the unified mathematical framework is designed, which shows that the intelligent transceiver based on the mathematical framework has higher performance.
Feng Li 0057, Yiqing Zhang 0001, Zhengyang Hu 0001, Guanzhang Liu, Runhua Li, Jiang Xue 0001, Zongben Xu
VTC Fall9
2022 On the uniqueness of virtual substrate for metasurface in a dielectric half-space
Xiaoming Chen 0002, Anxue Zhang, Qiang Cheng 0002, Zongben Xu
Sci. China Inf. Sci.6
2022 Unlabeled data driven cost-sensitive inverse projection sparse representation-based classification with 1/2 regularization
Jian Sun 0009, Zongben Xu
Sci. China Inf. Sci.4
2022 Dynamic Neural Network for MIMO Detection
abstract
Achieving adequate precision in deep learning based communications often requires large network architectures, which results into unacceptable time delay and power consumption. This paper introduces the dynamic neural network (DyNN) into the design of wireless communications systems. DyNN allocates different samples with computation resources on demand by preforming dynamic inferences, thereby reducing the redundant computational cost and enhancing the network efficiency. We design a dynamic depth architecture that allows samples to adaptively skip layers with various dynamic strategies, from which we further develop aconfidence criterion baseddynamicimproved DetNet (CD-IDetNet) and apolicy network baseddynamicimproved DetNet (PD-IDetNet) for multiple-input multiple-output (MIMO) detection. Specifically, in CD-IDetNet, a confidence criterion is adopted to control samples exiting early, while in PD-IDetNet, policy networks are trained by reinforcement learning to selectively skip layers for varying samples. Simulation results demonstrate that CD-IDetNet and PD-IDetNet detectors can respectively reduce 17.4% and 31.1% computational costs while preserving the full accuracy of IDetNet. Desirable tradeoffs between accuracy and computational complexity can be further achieved by fine-tuning the hyper-parameters of CD-IDetNet and PD-IDetNet. Moreover, over-the-air (OTA) tests are conducted to validate the effectiveness of the proposed detectors in practical systems.
Yuwen Yang, Feifei Gao 0001, Mingjin Wang, Jiang Xue 0001, Zongben Xu
IEEE J. Sel. Areas Commun.5
2022 Deep alternating non-negative matrix factorisation
Jianyong Sun, Qingming Kong, Zongben Xu
Knowl. Based Syst.3
2022 Variational HyperAdam: A Meta-Learning Approach to Network Training
abstract
Stochastic optimization algorithms have been popular for training deep neural networks. Recently, there emerges a new approach of learning-based optimizer, which has achieved promising performance for training neural networks. However, these black-box learning-based optimizers do not fully take advantage of the experience in human-designed optimizers and heavily rely on learning from meta-training tasks, therefore have limited generalization ability. In this paper, we propose a novel optimizer, dubbed as Variational HyperAdam, which is based on a parametric generalized Adam algorithm, i.e., HyperAdam, in a variational framework. With Variational HyperAdam as optimizer for training neural network, the parameter update vector of the neural network at each training step is considered as random variable, whose approximate posterior distribution given the training data and current network parameter vector is predicted by Variational HyperAdam. The parameter update vector for network training is sampled from this approximate posterior distribution. Specifically, in Variational HyperAdam, we design a learnable generalized Adam algorithm for estimating expectation, paired with a VarBlock for estimating the variance of the approximate posterior distribution of parameter update vector. The Variational HyperAdam is learned in a meta-learning approach with meta-training loss derived by variational inference. Experiments verify that the learned Variational HyperAdam achieved state-of-the-art network training performance for various types of networks on different datasets, such as multilayer perceptron, CNN, LSTM and ResNet.
Shipeng Wang 0002, Yan Yang 0007, Jian Sun 0009, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 MHF-Net: An Interpretable Deep Network for Multispectral and Hyperspectral Image Fusion
abstract
Multispectral and hyperspectral image fusion (MS/HS fusion) aims to fuse a high-resolution multispectral (HrMS) and a low-resolution hyperspectral (LrHS) images to generate a high-resolution hyperspectral (HrHS) image, which has become one of the most commonly addressed problems for hyperspectral image processing. In this paper, we specifically designed a network architecture for the MS/HS fusion task, called MHF-net, which not only contains clear interpretability, but also reasonably embeds the well studied linear mapping that links the HrHS image to HrMS and LrHS images. In particular, we first construct an MS/HS fusion model which merges the generalization models of low-resolution images and the low-rankness prior knowledge of HrHS image into a concise formulation, and then we build the proposed network by unfolding the proximal gradient algorithm for solving the proposed model. As a result of the careful design for the model and algorithm, all the fundamental modules in MHF-net have clear physical meanings and are thus easily interpretable. This not only greatly facilitates an easy intuitive observation and analysis on what happens inside the network, but also leads to its good generalization capability. Based on the architecture of MHF-net, we further design two deep learning regimes for two general cases in practice: consistent MHF-net and blind MHF-net. The former is suitable in the case that spectral and spatial responses of training and testing data are consistent, just as considered in most of the pervious general supervised MS/HS fusion researches. The latter ensures a good generalization in mismatch cases of spectral and spatial responses in training and testing data, and even across different sensors, which is generally considered to be a challenging issue for general supervised MS/HS fusion methods. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research.
Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Graph Neural Network Encoding for Community Detection in Attribute Networks
abstract
In this article, we first propose a graph neural network encoding method for the multiobjective evolutionary algorithm (MOEA) to handle the community detection problem in complex attribute networks. In the graph neural network encoding method, each edge in an attribute network is associated with a continuous variable. Through nonlinear transformation, a continuous valued vector (i.e., a concatenation of the continuous variables associated with the edges) is transferred to a discrete valued community grouping solution. Further, two objective functions for the single-attribute and multiattribute network are proposed to evaluate the attribute homogeneity of the nodes in communities, respectively. Based on the new encoding method and the two objectives, a MOEA based upon NSGA-II, called continuous encoding MOEA, is developed for the transformed community detection problem with continuous decision variables. Experimental results on single-attribute and multiattribute networks with different types show that the developed algorithm performs significantly better than some well-known evolutionary- and nonevolutionary-based algorithms. The fitness landscape analysis verifies that the transformed community detection problems have smoother landscapes than those of the original problems, which justifies the effectiveness of the proposed graph neural network encoding method.
Jianyong Sun, Wei Zheng 0004, Qingfu Zhang 0001, Zongben Xu
IEEE Trans. Cybern.4
2022 PanCSC-Net: A Model-Driven Deep Unfolding Method for Pansharpening
abstract
Recently, deep learning (DL) approaches have been widely applied to the pansharpening problem, which is defined as fusing a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image to obtain a high-resolution multispectral (HRMS) image. However, most DL-based methods handle this task by designing black-box network architectures to model the mapping relationship from LRMS and PAN to HRMS. These network architectures always lack sufficient interpretability, which limits their further performance improvements. To address this issue, we adopt the model-driven method to design an interpretable deep network structure for pansharpening. First, we present a new pansharpening model using the convolutional sparse coding (CSC), which is quite different from the current pansharpening frameworks. Second, an alternative algorithm is developed to optimize this model. This algorithm is further unfolded to a network, where each network module corresponds to a specific operation of the iterative algorithm. Therefore, the proposed network has clear physical interpretations, and all the learnable modules can be automatically learned in an end-to-end way from the given dataset. Experimental results on some benchmark datasets show that our network performs better than other advanced methods both quantitatively and qualitatively.
Xiangyong Cao, Xueyang Fu, Danfeng Hong, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.4
2022 A Deep-Learning-Based Generalized Convolutional Model For Seismic Data and Its Application in Seismic Deconvolution
abstract
The convolutional model, which describes the relation among poststack seismic data, wavelet, and reflectivity, is the foundation of seismic deconvolution (SD). However, this model is only an approximation of the seismic wave equation, and it may not work in complex cases especially when the medium is anelastic, heterogeneous, and anisotropic. In this article, we propose a generalized convolutional model for poststack seismic data. A deep-learning-based data correction term is added to characterize the data ingredients that cannot be characterized by the convolutional model. The data correction term of the new model is realized using the long-short term memory (LSTM)-based deep learning architecture, of which parameters are learned based on the dataset from several well logs. Based on the new model, we propose an SD method and investigate its performance in building reflectivity models using complex numerical examples. The results verified that the new model can accurately characterize complex seismic data, which cannot be characterized by a convolutional model. In addition, the proposed SD method has significant advantages over traditional methods in building high-fidelity reflectivity models in complex cases.
Zhaoqi Gao, Sichao Hu, Chuang Li 0003, Hongling Chen, Xiudi Jiang, Zhibin Pan, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.8
2022 Self-Supervised Deep Learning for Nonlinear Seismic Full Waveform Inversion
abstract
Seismic full waveform inversion (FWI) is able to build high-resolution velocity model based on the full information carried by seismic wave. However, FWI requires an accurate enough initial model to ensure convergence. In this paper, we propose a new nonlinear FWI method to mitigate the initial model dependence problem. Specifically, we firstly propose a nonlinear operator within the hybrid model- and data-driven framework based on the frequency controllable envelope operator (FCEO) and a deep learning architecture U-Net. FCEO is used to obtain the envelope of a band-limited data and U-Net realizes the mapping from this envelope to that corresponding to a lower frequency band. The U-Net is trained in a self-supervised manner that avoids the reliance on labeled data and benefits the generalization ability. Based on the nonlinear operator, a nonlinear FWI method is proposed by defining a new misfit function. In addition, the calculation of gradient is derived using the adjoint-state method. Using numerical examples, we investigate the performance of the proposed nonlinear operator and the new nonlinear FWI method. The results clearly demonstrate that the proposed nonlinear operator is effective in obtaining low-frequency envelope data, and the new nonlinear FWI method has advantages over common method in mitigating cycle-skipping and building an initial model for conventional FWI.
Zhaoqi Gao, Chuang Li 0003, Feipeng Li, Qingzhen Wang, Jicai Ding, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.8
2022 Quantum-Enhanced Deep Learning-Based Lithology Interpretation From Well Logs
abstract
Lithology interpretation is important for understanding subsurface properties. Yet, the common manual well log interpretation is usually with low efficiency and bad consistency. Therefore, the automatic well log interpretation tools based on machine learning and deep learning have been developed. Although the state-of-the-art sophisticated models can show fine interpretation performance with acceptable accuracies, “blind” tests do not always exhibit satisfactory results because of the complexity of lithology interpretation with respect to subsurface rock properties and the data-labeling quality. To solve this generalization challenge, we propose to leverage the parameterized quantum circuits in the deep-learning model. The quantum computing takes advantages of the superposition and entanglement quantum systems, which could potentially endow the generalization power or capability to the deep-learning model. Using the proposed quantum-enhanced deep-learning (QEDL) model, we have tested the model performance on field well log data from different wells. Compared with the classic fine convolutional neural network (CNN) model and the long short-term memory (LSTM) model, the proposed QEDL model achieves comparable model performance with a clearly improved generalization power for interpreting both thin and thick lithology layers. In addition, because of the quantum circuit structure, the QEDL model needs much fewer model parameters than LSTM and CNN models, i.e., the QEDL parameter number in our study can be approximately 75% less than that of LSTM and 89% less than that of CNN.
Naihao Liu, Jinghuai Gao, Zongben Xu, Daxing Wang, Fangyu Li 0002
IEEE Trans. Geosci. Remote. Sens.4
2022 Ground-Roll Separation and Attenuation Using Curvelet-Based Multichannel Variational Mode Decomposition
abstract
Ground roll, a source-generated surface wave, is a main type of coherent noise in a land seismic survey. Low-frequency and high-amplitude ground roll often overlays valid reflection events, resulting in obscuring seismic reflections. Ground-roll attenuation is an essential step for seismic data processing, which is based on the accurate separation of the ground roll and reflections without damaging their morphological characteristics. In this study, we effectively separate and suppress ground roll in shot gathers through a proposed workflow. We first identify the major components of the ground roll adopting the multichannel variational mode decomposition (MVMD), which shows significant improvements compared to the conventional single-channel VMD. Ground roll can be identified on the decomposed band-limited intrinsic mode functions (IMFs). Moreover, we propose an adaptive criterion to determine the number of decomposed IMFs. Due to the narrowly concentrated frequency components with the multichannel continuity constraint from MVMD, ground roll is mainly contained in low-frequency IMFs, which benefits the accurate ground-roll suppression. Next, we separate ground roll and reflections on the selected low-frequency IMFs through a curvelet based block-coordinate relaxation method. Afterward, we can obtain a filtered gather by removing the separated ground roll from the original shot gather. Finally, we apply the proposed workflow to synthetic and field gathers to testify its validity and effectiveness for simultaneously attenuating ground roll and preserving valid seismic reflector information.
Naihao Liu, Fangyu Li 0002, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.5
2022 Simultaneous Inversion for Reflectivity and Q Using Nonstationary Seismic Data With Deep-Learning-Based Decoupling
abstract
Building reflectivity and quality factor (Q) using nonstationary post-stack seismic data is important for vertical resolution enhancement of seismic data and reservoir identification. However, it is well-known that both reflectivity and Q affect the waveform of seismic data, leading to the fact that simultaneously estimating them is a strong ill-posed multi-parameter inverse problem which faces the crosstalk problem. In this paper, we propose a new method for simultaneous inversion of reflectivity and Q. A deep-learning-based data decoupling operator is proposed to decouple the effects of the two parameters on nonstationary seismic data. Based on the decoupled data, we transform the original multi-parameter inverse problem into two independent singe-parameter inverse problems that are immune to crosstalk and can build reasonable initial models for reflectivity and Q. Then alternative iteration is conducted to update the two built initial models to obtain the final models. A few well-logs are used to train the deep learning architecture and specific regularization terms are constructed for the inverse problem to ensure physically reasonable results. Synthetic and field data examples verify the effectiveness of the proposed method and its advantages over a conventional model-driven joint inversion method.
Linan Xu, Zhaoqi Gao, Sichao Hu, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.5
2022 Semi-Active Convolutional Neural Networks for Hyperspectral Image Classification
abstract
Owing to the powerful data representation ability of deep learning (DL) techniques, tremendous progress has been recently made in hyperspectral image (HSI) classification. Convolutional neural network (CNN), as a main part of the DL family, has been proven to be considerably effective to extract spatial-spectral features for HSIs. Nevertheless, its classification performance, to a great extent, depends on the quality and quantity of samples in the network training process. To select those samples, either labeled or unlabeled, that can be used to enhance the generalization ability of CNNs and further improve the classification accuracy, we propose an iterative semi-supervised CNNs framework by means of active learning and superpixel segmentation techniques, dubbed as semi-active CNNs (SA-CNNs), for HSI classification. More specifically, we start to pre-train a CNNs-based model on a small-scale unbiased labeled set and infer unlabeled data using the trained model, i.e., generating pseudo-labels. Then, the reliable samples, which consist of two parts: high label-homogeneity and most informativeness, are actively selected from superpixel segments. These selected labeled and unlabeled samples with their labels and pseudo-labels are re-fed into the next-round network training. Moreover, three different schedules, i.e.,log-,exp-, andlinear-schedules, are progressively adopted to fully explore their potentials in sample selection, until a labeling budget is finally reached. Extensive experiments are conducted on three benchmark HSI datasets, demonstrating substantial performance improvements of the proposed SA-CNNs over other similar competitors.
Jing Yao 0002, Xiangyong Cao, Danfeng Hong, Xin Wu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.7
2022 Sparsity-Enhanced Convolutional Decomposition: A Novel Tensor-Based Paradigm for Blind Hyperspectral Unmixing
abstract
Blind hyperspectral unmixing (HU) has long been recognized as a crucial component in analyzing the hyperspectral imagery (HSI) collected by airborne and spaceborne sensors. Due to the highly ill-posed problems of such a blind source separation scheme and the effects of spectral variability in hyperspectral imaging, the ability to accurately and effectively unmixing the complex HSI still remains limited. To this end, this article presents a novel blind HU model, called sparsity-enhanced convolutional decomposition (SeCoDe), by jointly capturing spatial–spectral information of HSI in a tensor-based fashion. SeCoDe benefits from two perspectives. On the one hand, the convolutional operation is employed in SeCoDe to locally model the spatial relation between the targeted pixel and its neighbors, which can be well explained by spectral bundles that are capable of addressing spectral variabilities effectively. It maintains, on the other hand, physically continuous spectral components by decomposing the HSI along with the spectral domain. With sparsity-enhanced regularization, an alternative optimization strategy with alternating direction method of multipliers (ADMM)-based optimization algorithm is devised for efficient model inference. Extensive experiments conducted on three different data sets demonstrate the superiority of the proposed SeCoDe compared to previous state-of-the-art methods. We will also release the code athttps://github.com/danfenghong/IEEE_TGRS_SeCoDeto encourage the reproduction of the given results.
Jing Yao 0002, Danfeng Hong, Lin Xu 0001, Deyu Meng, Jocelyn Chanussot, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.6
2022 Robust Online CSI Estimation in a Complex Environment
abstract
Channel state information (CSI) estimation is one of the key techniques for improving the performance of wireless communication systems. Meanwhile, the fifth generation wireless communication systems require higher accuracy and lower latency for CSI estimation. In this paper, the methods of noise modeling and online learning are combined to improve the accuracy and reduce the latency. The complex noise environment (considering noise and interference together) is modeled as a specific mixture of Gaussian (MoG) distribution because of its widely approximation capability to any continuous distribution. The MoG CSI estimation (MoG-CE) model and expectation maximization (EM) algorithm are introduced as one of the baseline methods. Further, the parameters of the model can be updated in real time based on the prior knowledge of historical information. Therefore, the online MoG CSI estimation (O-MoG-CE) model and online MoG dynamic CSI estimation (O-MoG-D-CE) model are proposed for time-invariant and time-varying CSI estimations, respectively. The above models can not only self-adapt to various complex communication scenarios robustly but also achieve online and dynamic CSI estimation to improve the accuracy and reduce the latency significantly. In addition, the proposed models can be formulated as standard maximum a posteriori estimations and efficient online expectation maximization (OEM) algorithms are applied for the estimations in a pure machine learning fashion. Comparing with baseline methods, the simulation results demonstrate the superiority of the proposed methods in terms of the accuracy, latency and computation consumption.
Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu
IEEE Trans. Wirel. Commun.6
2021 Learning to Mutate for Differential Evolution
abstract
Adaptive parameter control and mutation operator selection are two important research avenues in differential evolution (DE). Existing works consider the two avenues independently. In this paper, we propose to unify the two modules and develop a unified parameterized mutation operator. With different settings of the parameters, different mutation operators can be retrieved. Further, the settings of the parameters closely relate to the control parameters of the DE. By determining the parameters we can achieve adaptive parameter control and mutation operator selection simultaneously. We propose to use a neural network to output the parameters and learn the network parameter by the natural evolution strategies algorithm under the consideration of modeling the evolution process as a Markov Decision Process. Experimental results on the CEC 2018 test suite show that the proposed method performs significantly better than traditional DEs with different operators and an advanced adaptive DE. We further analyze the time complexity and population diversity of the proposed method. The analysis shows that our method can achieve a balanced exploration and exploitation with a properly learned network.
Haotian Zhang 0023, Jianyong Sun, Zongben Xu
CEC3
2021 Training Networks in Null Space of Feature Covariance for Continual Learning
abstract
In the setting of continual learning, a network is trained on a sequence of tasks, and suffers from catastrophic forgetting. To balance plasticity and stability of network in continual learning, in this paper, we propose a novel network training algorithm Adam-NSCL which sequentially optimizes network parameters in the null space of all previous tasks. We first propose two mathematical conditions respectively for achieving network stability and plasticity in continual learning. Based on them, the network training for sequential tasks without forgetting can be simply achieved by projecting the candidate parameter update into the approximate null space of all previous tasks in the network training process, where the candidate parameter update can be generated by Adam. The approximate null space can be derived by applying singular value decomposition to the un-centered covariance matrix of all input features of previous tasks for each linear layer. For efficiency, the uncentered covariance matrix can be incrementally computed after learning each task. We also empirically verify the rationality of the approximate null space at each linear layer. We apply our approach to training networks for continual learning on benchmark datasets of CIFAR-100 and TinyImageNet, and the results suggest that the proposed approach outperforms or matches the state-ot-the-art continual learning approaches.
Shipeng Wang 0002, Xiaorong Li, Jian Sun 0009, Zongben Xu
CVPR4
2021 A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation
Heran Yang, Jian Sun 0009, Zongben Xu
MICCAI (3)4
2021 Adversarial Reweighting for Partial Domain Adaptation
abstract
Partial domain adaptation (PDA) has gained much attention due to its practical setting. The current PDA methods usually adapt the feature extractor by aligning the target and reweighted source domain distributions. In this paper, we experimentally find that the feature adaptation by the reweighted distribution alignment in some state-of-the-art PDA methods is not robust to the ``noisy'' weights of source domain data, leading to negative domain transfer on some challenging benchmarks. To tackle the challenge of negative domain transfer, we propose a novel Adversarial Reweighting (AR) approach that adversarially learns the weights of source domain data to align the source and target domain distributions, and the transferable deep recognition network is learned on the reweighted source domain data. Based on this idea, we propose a training algorithm that alternately updates the parameters of the network and optimizes the weights of source domain data. Extensive experiments show that our method achieves state-of-the-art results on the benchmarks of ImageNet-Caltech, Office-Home, VisDA-2017, and DomainNet. Ablation studies also confirm the effectiveness of our approach.
Xiang Gu 0005, Yan Yang 0007, Jian Sun 0009, Zongben Xu
NeurIPS5
2021 A distribution independence based method for 3D face shape decomposition
Cuican Yu, Huibin Li 0001, Jian Sun 0009, Zongben Xu
Comput. Vis. Image Underst.5
2021 Robust online rain removal for surveillance videos with dynamic rains
Lixuan Yi, Qian Zhao 0002, Wei Wei 0006, Zongben Xu
Knowl. Based Syst.4
2021 OMMDE-Net: A Deep Learning-Based Global Optimization Method for Seismic Inversion
abstract
In this letter, we propose a new global optimization method for nonlinear seismic inversion problems. The proposed method is a development of the existing method MMDE-Net by introducing a learnable strategy for choosing problem-dependent basis vectors and regularization parameters that are considered to be fixed in MMDE-Net. We name the proposed method as the optimized MMDE-Net (OMMDE-Net) and investigate its performance in seismic inversion through both synthetic and field data examples. The experimental results demonstrate that OMMDE-Net has advantages over MMDE-Net in effectiveness and efficiency.
Zhaoqi Gao, Chuang Li 0003, Zhibin Pan, Jinghuai Gao, Zongben Xu
IEEE Geosci. Remote. Sens. Lett.6
2021 Sparse deep dictionary learning identifies differences of time-varying functional connectivity in brain neuro-developmental study
Chen Qiao, Lan Yang 0010, Vince D. Calhoun, Zongben Xu, Yu-Ping Wang 0002
Neural Networks4
2021 SPLBoost: An Improved Robust Boosting Algorithm Based on Self-Paced Learning
abstract
It is known that boosting can be interpreted as an optimization technique to minimize an underlying loss function. Specifically, the underlying loss being minimized by the traditional AdaBoost is the exponential loss, which proves to be very sensitive to random noise/outliers. Therefore, several boosting algorithms, e.g., LogitBoost and SavageBoost, have been proposed to improve the robustness of AdaBoost by replacing the exponential loss with some designed robust loss functions. In this article, we present a new way to robustify AdaBoost, that is, incorporating the robust learning idea of self-paced learning (SPL) into the boosting framework. Specifically, we design a new robust boosting algorithm based on the SPL regime, that is, SPLBoost, which can be easily implemented by slightly modifying off-the-shelf boosting packages. Extensive experiments and a theoretical characterization are also carried out to illustrate the merits of the proposed SPLBoost.
Kaidong Wang, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Xiuwu Liao, Zongben Xu
IEEE Trans. Cybern.6
2021 Learning Adaptive Differential Evolution Algorithm From Optimization Experiences by Policy Gradient
abstract
Differential evolution is one of the most prestigious population-based stochastic optimization algorithm for black-box problems. The performance of a differential evolution algorithm depends highly on its mutation and crossover strategy and associated control parameters. However, the determination process for the most suitable parameter setting is troublesome and time consuming. Adaptive control parameter methods that can adapt to problem landscape and optimization environment are more preferable than fixed parameter settings. This article proposes a novel adaptive parameter control approach based on learning from the optimization experiences over a set of problems. In the approach, the parameter control is modeled as a finite-horizon Markov decision process. A reinforcement learning algorithm, named policy gradient, is applied to learn an agent (i.e., parameter controller) that can provide the control parameters of a proposed differential evolution adaptively during the search procedure. The differential evolution algorithm based on the learned agent is compared against nine well-known evolutionary algorithms on the CEC'13 and CEC'17 test suites. Experimental results show that the proposed algorithm performs competitively against these compared algorithms on the test suites.
Jianyong Sun, Xin Liu 0078, Thomas Bäck, Zongben Xu
IEEE Trans. Evol. Comput.4
2021 Large-Dimensional Seismic Inversion Using Global Optimization With Autoencoder-Based Model Dimensionality Reduction
abstract
Seismic inversion problems often involve strong nonlinear relationships between model and data so that their misfit functions usually have many local minima. Global optimization methods are well known to be able to find the global minimum without requiring an accurate initial model. However, when the dimensionality of model space becomes large, global optimization methods will converge slow, which seriously hinders their applications in large-dimensional seismic inversion problems. In this article, we propose a new method for large-dimensional seismic inversion based on global optimization and a machine learning technique called autoencoder. Benefiting from the dimensionality reduction characteristics of autoencoder, the proposed method converts the original large-dimensional seismic inversion problem into a low-dimensional one that can be effectively and efficiently solved by global optimization. We apply the proposed method to seismic impedance inversion problems to test its performance. We use a trace-by-trace inversion strategy, and regularization is used to guarantee the lateral continuity of the inverted model. Well-log data with accurate velocity and density are the prerequisite of the inversion strategy to work effectively. Numerical results of both synthetic and field data examples clearly demonstrate that the proposed method can converge faster and yield better inversion results compared with common methods.
Zhaoqi Gao, Chuang Li 0003, Naihao Liu, Zhibin Pan, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.6
2021 Multimodal GANs: Toward Crossmodal Hyperspectral-Multispectral Image Segmentation
abstract
This article addresses the problem of semantic segmentation with limited cross-modality data in large-scale urban scenes. Most prior works have attempted to address this issue by using multimodal deep neural networks (DNNs). However, their ability to effectively blending different properties across multimodalities and robustly learning representations from complex scenes remains limited, particularly in the absence of sufficient and well-annotated training images. This leads to a challenge related to cross-modality learning with multimodal DNNs. To this end, we introduce two novel plug-and-play units in the network: self-generative adversarial networks (GANs) module and mutual-GANs module, to learn perturbation-insensitive feature representations and to eliminate the gap between multimodalities, respectively, yielding more effective and robust information transfer. Furthermore, a patchwise progressive training strategy is devised to enable effective network learning with limited samples. We evaluate the proposed network on two multimodal (hyperspectral and multispectral) overhead image data sets and achieve a significant improvement in comparison with several state-of-the-art methods.
Danfeng Hong, Jing Yao 0002, Deyu Meng, Zongben Xu, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2020 Adaptive Structural Hyper-Parameter Configuration by Q-Learning
abstract
Tuning hyper-parameters for evolutionary algorithms is an important issue in computational intelligence. Performance of an evolutionary algorithm depends not only on its operation strategy design, but also on its hyper-parameters. Hyper-parameters can be categorized in two dimensions as structural/numerical and time-invariant/time-variant. Particularly, structural hyper-parameters in existing studies are usually tuned in advance for time-invariant parameters, or with hand-crafted scheduling for time-invariant parameters. In this paper, we make the first attempt to model the tuning of structural hyper-parameters as a reinforcement learning problem, and present to tune the structural hyper-parameter which controls computational resource allocation in the CEC 2018 winner algorithm by Q-learning. Experimental results show favorably against the winner algorithm on the CEC 2018 test functions.
Haotian Zhang 0023, Jianyong Sun, Zongben Xu
CEC3
2020 Spherical Space Domain Adaptation With Robust Pseudo-Label Loss
abstract
Adversarial domain adaptation (DA) has been an effective approach for learning domain-invariant features by adversarial training. In this paper, we propose a novel adversarial DA approach completely defined in spherical feature space, in which we define spherical classifier for label prediction and spherical domain discriminator for discriminating domain labels. To utilize pseudo-label robustly, we develop a robust pseudo-label loss in the spherical feature space, which weights the importance of estimated labels of target data by posterior probability of correct labeling, modeled by Gaussian-uniform mixture model in spherical feature space. Extensive experiments show that our method achieves state-of-the-art results, and also confirm effectiveness of spherical classifier, spherical discriminator and spherical robust pseudo-label loss.
Xiang Gu 0005, Jian Sun 0009, Zongben Xu
CVPR3
2020 Cross-Attention in Coupled Unmixing Nets for Unsupervised Hyperspectral Super-Resolution
Jing Yao 0002, Danfeng Hong, Jocelyn Chanussot, Deyu Meng, Xiao Xiang Zhu 0001, Zongben Xu
ECCV (29)6
2020 MEP-Based Channel Estimation under Complex Communication Environment
abstract
In this paper, we study the channel state information (CSI) estimation by utilizing maximum entropy principle (MEP) and noise modeling method. The new model can not only represent the characters of the complex communication environment, but can also adjust itself according to the environment by using machine learning. In addition, a new iteration algorithm is presented to derive numerical results. Adaptive parameters learning and features choosing capability make the proposed method outperform the existing methods. The accuracy of estimation is verified by the Monte Carlo simulations.
Zhengyang Hu 0001, Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu
ICC5
2020 Model-Driven Deep Attention Network for Ultra-fast Compressive Sensing MRI Guided by Cross-contrast MR Image
Yan Yang 0007, Heran Yang, Jian Sun 0009, Zongben Xu
MICCAI (2)5
2020 On Presuppositions of Machine Learning: A Meta Theory
abstract
Machine learning (ML) has been run and applied by premising a series of presuppositions, which contributes both the great success of AI and the bottleneck of further development of ML. These presuppositions include (i) the independence assumption of loss function on dataset (Hypothesis I); (ii) the large capacity assumption on hypothesis space including solution (Hypothesis II); (iii) the completeness assumption of training data with high quality (Hypothesis III); and (iv) the Euclidean assumption on analysis framework and methodology (Hypothesis IV).
Zongben Xu
SIGIR1
2020 Color and direction-invariant nonlocal self-similarity prior and its application to color image denoising
Qi Xie 0002, Qian Zhao 0002, Zongben Xu, Deyu Meng
Sci. China Inf. Sci.3
2020 ADMM-CSNet: A Deep Learning Approach for Image Compressive Sensing
abstract
Compressive sensing (CS) is an effective technique for reconstructing image from a small amount of sampled data. It has been widely applied in medical imaging, remote sensing, image compression, etc. In this paper, we propose two versions of a novel deep learning architecture, dubbed as ADMM-CSNet, by combining the traditional model-based CS method and data-driven deep learning method for image reconstruction from sparsely sampled measurements. We first consider a generalized CS model for image reconstruction with undetermined regularizations in undetermined transform domains, and then two efficient solvers using Alternating Direction Method of Multipliers (ADMM) algorithm for optimizing the model are proposed. We further unroll and generalize the ADMM algorithm to be two deep architectures, in which all parameters of the CS model and the ADMM algorithm are discriminatively learned by end-to-end training. For both applications of fast CS complex-valued MR imaging and CS imaging of real-valued natural images, the proposed ADMM-CSNet achieved favorable reconstruction accuracy in fast computational speed compared with the traditional and the other deep learning methods.
Yan Yang 0007, Jian Sun 0009, Huibin Li 0001, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Abnormality detection in retinal image by individualized background learning
Benzhi Chen, Lisheng Wang, Xiuying Wang 0001, Jian Sun 0009, David Dagan Feng, Zongben Xu
Pattern Recognit.7
2020 Learning Through Deterministic Assignment of Hidden Parameters
abstract
Supervised learning frequently boils down to determining hidden and bright parameters in a parameterized hypothesis space based on finite input-output samples. The hidden parameters determine the nonlinear mechanism of an estimator, while the bright parameters characterize the linear mechanism. In a traditional learning paradigm, hidden and bright parameters are not distinguished and trained simultaneously in one learning process. Such a one-stage learning (OSL) brings a benefit of theoretical analysis but suffers from the high computational burden. In this paper, we propose a two-stage learning scheme, learning through deterministic assignment of hidden parameters (LtDaHPs), suggesting to deterministically generate the hidden parameters by using minimal Riesz energy points on a sphere and equally spaced points in an interval. We theoretically show that with such a deterministic assignment of hidden parameters, LtDaHP with a neural network realization almost shares the same generalization performance with that of OSL. Then, LtDaHP provides an effective way to overcome the high computational burden of OSL. We present a series of simulations and application examples to support the outperformance of LtDaHP.
Jian Fang 0001, Shaobo Lin, Zongben Xu
IEEE Trans. Cybern.3
2020 Hyperspectral Image Classification With Convolutional Neural Network and Active Learning
abstract
Deep neural network has been extensively applied to hyperspectral image (HSI) classification recently. However, its success is greatly attributed to numerous labeled samples, whose acquisition costs a large amount of time and money. In order to improve the classification performance while reducing the labeling cost, this article presents an active deep learning approach for HSI classification, which integrates both active learning and deep learning into a unified framework. First, we train a convolutional neural network (CNN) with a limited number of labeled pixels. Next, we actively select the most informative pixels from the candidate pool for labeling. Then, the CNN is fine-tuned with the new training set constructed by incorporating the newly labeled pixels. This step together with the previous step is iteratively conducted. Finally, Markov random field (MRF) is utilized to enforce class label smoothness to further boost the classification performance. Compared with the other state-of-the-art traditional and deep learning-based HSI classification methods, our proposed approach achieves better performance on three benchmark HSI data sets with significantly fewer labeled samples.
Xiangyong Cao, Jing Yao 0002, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.3
2020 Hyperspectral and Multispectral Image Fusion via Nonlocal Low-Rank Tensor Decomposition and Spectral Unmixing
abstract
Hyperspectral (HS) imaging has shown its superiority in many real applications. However, it is usually difficult to obtain high-resolution (HR) HS images through existing imaging techniques due to the hardware limitations. To improve the spatial resolution of HS images, this article proposes an effective HS-multispectral (HS-MS) image fusion method by combining the ideas of nonlocal low-rank tensor modeling and spectral unmixing. To be more precise, instead of unfolding the HS image into a matrix as done in the literature, we directly represent it as a tensor, then a designed nonlocal Tucker decomposition is used to model its underlying spatial-spectral correlation and the spatial self-similarity. The MS image serves mainly as a data constraint to maintain spatial consistency. To further reduce the spectral distortions in spatial enhancement, endmembers, and abundances from the spectral are used for spectral regularization. An efficient algorithm based on the alternating direction method of multipliers (ADMM) is developed to solve the resulting model. Extensive experiments on four HS image data sets demonstrate the superiority of the proposed method over several state-of-the-art HS-MS image fusion methods.
Kaidong Wang, Yao Wang 0003, Xi-Le Zhao, Jonathan Cheung-Wai Chan, Zongben Xu, Deyu Meng
IEEE Trans. Geosci. Remote. Sens.5
2020 Polarimetric SAR Image Semantic Segmentation With 3D Discrete Wavelet Transform and Markov Random Field
abstract
Polarimetric synthetic aperture radar (PolSAR) image segmentation is currently of great importance in image processing for remote sensing applications. However, it is a challenging task due to two main reasons. Firstly, the label information is difficult to acquire due to high annotation costs. Secondly, the speckle effect embedded in the PolSAR imaging process remarkably degrades the segmentation performance. To address these two issues, we present a contextual PolSAR image semantic segmentation method in this paper. With a newly defined channel-wise consistent feature set as input, the three-dimensional discrete wavelet transform (3D-DWT) technique is employed to extract discriminative multi-scale features that are robust to speckle noise. Then Markov random field (MRF) is further applied to enforce label smoothness spatially during segmentation. By simultaneously utilizing 3D-DWT features and MRF priors for the first time, contextual information is fully integrated during the segmentation to ensure accurate and smooth segmentation. To demonstrate the effectiveness of the proposed method, we conduct extensive experiments on three real benchmark PolSAR image data sets. Experimental results indicate that the proposed method achieves promising segmentation accuracy and preferable spatial consistency using a minimal number of labeled pixels.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Yong Xue, Zongben Xu
IEEE Trans. Image Process.5
2020 Unsupervised MR-to-CT Synthesis Using Structure-Constrained CycleGAN
abstract
Synthesizing a CT image from an available MR image has recently emerged as a key goal in radiotherapy treatment planning for cancer patients. CycleGANs have achieved promising results on unsupervised MR-to-CT image synthesis; however, because they have no direct constraints between input and synthetic images, cycleGANs do not guarantee structural consistency between these two images. This means that anatomical geometry can be shifted in the synthetic CT images, clearly a highly undesirable outcome in the given application. In this paper, we propose a structure-constrained cycleGAN for unsupervised MR-to-CT synthesis by defining an extra structure-consistency loss based on the modality independent neighborhood descriptor. We also utilize a spectral normalization technique to stabilize the training process and a self-attention module to model the long-range spatial dependencies in the synthetic images. Results on unpaired brain and abdomen MR-to-CT image synthesis show that our method produces better synthetic CT images in both accuracy and visual quality as compared to other unsupervised synthesis methods. We also show that an approximate affine pre-registration for unpaired training data can improve synthesis results.
Heran Yang, Jian Sun 0009, Aaron Carass, Can Zhao 0001, Jerry L. Prince, Zongben Xu
IEEE Trans. Medical Imaging7
2020 Full-Spectrum-Knowledge-Aware Tensor Model for Energy-Resolved CT Iterative Reconstruction
abstract
Energy-resolved computed tomography (ErCT) with a photon counting detector concurrently produces multiple CT images corresponding to different photon energy ranges. It has the potential to generate energy-dependent images with improved contrast-to-noise ratio and sufficient material-specific information. Since the number of detected photons in one energy bin in ErCT is smaller than that in conventional energy-integrating CT (EiCT), ErCT images are inherently more noisy than EiCT images, which leads to increased noise and bias in the subsequent material estimation. In this work, we first deeply analyze the intrinsic tensor properties of two-dimensional (2D) ErCT images acquired in different energy bins and then present a F ull- S pectrum-knowledge-aware Tensor analysis and processing (FSTensor) method for ErCT reconstruction to suppress noise-induced artifacts to obtain high-quality ErCT images and high-accuracy material images. The presented method is based on three considerations: (1) 2D ErCT images obtained in different energy bins can be treated as a 3-order tensor with three modes, i.e., width, height and energy bin, and a rich global correlation exists among the three modes, which can be characterized by tensor decomposition. (2) There is a locally piecewise smooth property in the 3-order ErCT images, and it can be captured by a tensor total variation regularization. (3) The images from the full spectrum are much better than the ErCT images with respect to noise variance and structural details and serve as external information to improve the reconstruction performance. We then develop an alternating direction method of multipliers algorithm to numerically solve the presented FSTensor method. We further utilize a genetic algorithm to tackle the parameter selection in ErCT reconstruction, instead of manually determining parameters. Simulation, preclinical and synthesized clinical ErCT results demonstrate that the presented FSTensor method leads to significant improvements over the filtered back-projection, robust principal component analysis, tensor-based dictionary learning and low-rank tensor decomposition with spatial-temporal total variation methods.
Dong Zeng, Yongshuai Ge, Sui Li, Qi Xie 0002, Hao Zhang 0026, Zhaoying Bian, Qian Zhao 0002, Yuanqing Li 0001, Zongben Xu, Deyu Meng, Jianhua Ma 0001
IEEE Trans. Medical Imaging10
2020 Learning to Search for MIMO Detection
abstract
This paper proposes a novel learning to learn method, called learning to learn iterative search algorithm (LISA), for signal detection in a multi-input multi-output (MIMO) system. The idea is to regard the signal detection problem as a decision making problem over tree. The goal is to learn the optimal decision policy. In LISA, deep neural networks are used as parameterized policy function. Through training, optimal parameters of the neural networks are learned and thus optimal policy can be approximated. Different neural network-based architectures are used for fixed and varying channel models, respectively. LISA provides soft decisions and does not require any information about the additive white Gaussian noise. Simulation results show that LISA 1) obtains near maximum likelihood detection performance in both fixed and varying channel models under QPSK modulation; 2) achieves significantly better bit error rate (BER) performance than classical detectors and recently proposed deep/machine learning based detectors at various modulations and signal to noise (SNR) ratios both under i.i.d and correlated Rayleigh fading channels in the simulation experiments; 3) is robust to MIMO detection problems with imperfect channel state information; and 4) generalizes very well against channel correlation and SNRs.
Jianyong Sun, Yiqing Zhang 0001, Jiang Xue 0001, Zongben Xu
IEEE Trans. Wirel. Commun.4
2019 HyperAdam: A Learnable Task-Adaptive Adam for Network Training
abstract
Deep neural networks are traditionally trained using humandesigned stochastic optimization algorithms, such as SGD and Adam. Recently, the approach of learning to optimize network parameters has emerged as a promising research topic. However, these learned black-box optimizers sometimes do not fully utilize the experience in human-designed optimizers, therefore have limitation in generalization ability. In this paper, a new optimizer, dubbed as HyperAdam, is proposed that combines the idea of “learning to optimize” and traditional Adam optimizer. Given a network for training, its parameter update in each iteration generated by HyperAdam is an adaptive combination of multiple updates generated by Adam with varying decay rates . The combination weights and decay rates in HyperAdam are adaptively learned depending on the task. HyperAdam is modeled as a recurrent neural network with AdamCell, WeightCell and StateCell. It is justified to be state-of-the-art for various network training, such as multilayer perceptron, CNN and LSTM.
Shipeng Wang 0002, Jian Sun 0009, Zongben Xu
AAAI3
2019 Semi-Supervised Transfer Learning for Image Rain Removal
abstract
Single image rain removal is a typical inverse problem in computer vision. The deep learning technique has been verified to be effective for this task and achieved state-of-the-art performance. However, previous deep learning methods need to pre-collect a large set of image pairs with/without synthesized rain for training, which tends to make the neural network be biased toward learning the specific patterns of the synthesized rain, while be less able to generalize to real test samples whose rain types differ from those in the training data. To this issue, this paper firstly proposes a semi-supervised learning paradigm toward this task. Different from traditional deep learning methods which only use supervised image pairs with/without synthesized rains, we further put real rainy images, without need of their clean ones, into the network training process. This is realized by elaborately formulating the residual between an input rainy image and its expected network output (clear image without rain) as a concise mixture of Gaussians distribution. The network is therefore trained to transfer to adapting the real rain pattern domain instead of only the synthesis rain domain, and thus both the short-of-training-sample and bias-to-supervised-sample issues can be evidently alleviated. Experiments on synthetic and real data verify the superiority of our model compared to the state-of-the-arts.
Wei Wei 0006, Deyu Meng, Qian Zhao 0002, Zongben Xu
CVPR4
2019 Multispectral and Hyperspectral Image Fusion by MS/HS Fusion Net
abstract
Hyperspectral imaging can help better understand the characteristics of different materials, compared with traditional image systems. However, only high-resolution multispectral (HrMS) and low-resolution hyperspectral (LrHS) images can generally be captured at video rate in practice. In this paper, we propose a model-based deep learning approach for merging an HrMS and LrHS images to generate a high-resolution hyperspectral (HrHS) image. In specific, we construct a novel MS/HS fusion model which takes the observation models of low-resolution images and the low-rankness knowledge along the spectral mode of HrHS image into consideration. Then we design an iterative algorithm to solve the model by exploiting the proximal gradient method. And then, by unfolding the designed algorithm, we construct a deep network, called MS/HS Fusion Net, with learning the proximal operators and model parameters by convolutional neural networks. Experimental results on simulated and real data substantiate the superiority of our method both visually and quantitatively as compared with state-of-the-art methods along this line of research.
Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Wangmeng Zuo, Zongben Xu
CVPR6
2019 Robust CSI Estimation Under Complex Communication Environment
abstract
Channel estimation is the critical and fundamental problem in wireless communication techniques, however, the complexity environment, including interference and noise, post a fundamental limit on the accuracy of channel estimation on practical applications. Most existing channel estimation techniques are based on the simple assumption of Gaussian white noise, which makes the performance poorly within real communication environment. To address this problem, we propose a new channel estimation method by assuming the environment as Mixture of Gaussian (MoG) distributions and penalized MoG (PMoG) model by combining the penalized likelihood method with MoG distributions. This model is proposed by the first time in the research of wireless communication, and the superiority of this method lies on its approximation capability to wide range of scenarios of complex communication environments adaptively and analyzing the environment by learning the proper number of statistical components. Moreover, we design an Expectation Maximization (EM) algorithm to estimate the parameters of the PMoG model. The advantage of our method is demonstrated by simulation experiments.
Haipei Zhang, Jiang Xue 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu
ICC5
2019 Unsupervised PolSAR Image Factorization with Deep Convolutional Networks
abstract
This paper presents a novel unsupervised polarimetric synthetic aperture radar (PolSAR) image classification method, which incorporates polarimetric image factorization and deep convolutional networks into a principled framework. To implement this idea, we design a convolutional neural network (CNN) with a newly defined loss function which measures the probability distribution distance between the initial distribution maps and CNN predictions. In the proposed method, we firstly execute polarimetric image factorization to generate a dictionary of meaningful atom scatters and their corresponding distribution maps, where the strongest scatters are selected as training samples for CNN. Next, we train the CNN by iteratively optimizing the defined energy function, producing the final distribution maps and classification result. The proposed approach is applied on a real UAVSAR image. Experimental results justify that our approach can effectively classify the PolSAR image in an unsupervised way and produce favorable classification results.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yibo Han, Yuanlong Cui, Yong Xue, Zongben Xu
IGARSS7
2019 An Active Deep Learning Approach for Minimally-Supervised Polsar Image Classification
abstract
Aiming at improving the classification performance with greatly reduced annotation cost, this paper presents an active deep learning approach for minimally-supervised PolSAR image classification, which integrates active learning and fine-tuning convolutional neural network (CNN) into a principled framework. Starting from a CNN trained using a very limited number of labeled pixels, we iteratively and actively select the most informative candidates for annotation, and incrementally fine-tune the CNN by incorporating the newly annotated pixels. Moreover, to boost the performance and robustness of the proposed method, we employ Markov random field to enforce label smoothness, and data augmentation technique to enlarge the training set. Extensive experiments demonstrated that our approach achieved state-of-the-art classification results with significantly reduced annotation cost.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yibo Han, Yuanlong Cui, Yong Xue, Zongben Xu
IGARSS7
2019 A Prior Learning Network for Joint Image and Sensitivity Estimation in Parallel MR Imaging
Nan Meng, Yan Yang 0007, Zongben Xu, Jian Sun 0009
MICCAI (4)3
2019 Neural Diffusion Distance for Image Segmentation
abstract
Diffusion distance is a spectral method for measuring distance among nodes on graph considering global data structure. In this work, we propose a spec-diff-net for computing diffusion distance on graph based on approximate spectral decomposition. The network is a differentiable deep architecture consisting of feature extraction and diffusion distance modules for computing diffusion distance on image by end-to-end training. We design low resolution kernel matching loss and high resolution segment matching loss to enforce the network's output to be consistent with human-labeled image segments. To compute high-resolution diffusion distance or segmentation mask, we design an up-sampling strategy by feature-attentional interpolation which can be learned when training spec-diff-net. With the learned diffusion distance, we propose a hierarchical image segmentation method outperforming previous segmentation methods. Moreover, a weakly supervised semantic segmentation network is designed using diffusion distance and achieved promising results on PASCAL VOC 2012 segmentation dataset.
Jian Sun 0009, Zongben Xu
NeurIPS2
2019 Meta-Weight-Net: Learning an Explicit Mapping For Sample Weighting
abstract
Current deep neural networks(DNNs) can easily overfit to biased training data with corrupted labels or class imbalance. Sample re-weighting strategy is commonly used to alleviate this issue by designing a weighting function mapping from training loss to sample weight, and then iterating between weight recalculating and classifier updating. Current approaches, however, need manually pre-specify the weighting function as well as its additional hyper-parameters. It makes them fairly hard to be generally applied in practice due to the significant variation of proper weighting schemes relying on the investigated problem and training data. To address this issue, we propose a method capable of adaptively learning an explicit weighting function directly from data. The weighting function is an MLP with one hidden layer, constituting a universal approximator to almost any continuous functions, making the method able to fit a wide range of weighting function forms including those assumed in conventional research. Guided by a small amount of unbiased meta-data, the parameters of the weighting function can be finely updated simultaneously with the learning process of the classifiers. Synthetic and real experiments substantiate the capability of our method for achieving proper weighting functions in class imbalance and noisy label cases, fully complying with the common settings in traditional methods, and more complicated scenarios beyond conventional cases. This naturally leads to its better accuracy than other state-of-the-art methods.
Qi Xie 0002, Lixuan Yi, Qian Zhao 0002, Sanping Zhou, Zongben Xu, Deyu Meng
NeurIPS6
2019 Joint analysis of individual-level and summary-level GWAS data by leveraging pleiotropy
abstract
MOTIVATION: A large number of recent genome-wide association studies (GWASs) for complex phenotypes confirm the early conjecture for polygenicity, suggesting the presence of large number of variants with only tiny or moderate effects. However, due to the limited sample size of a single GWAS, many associated genetic variants are too weak to achieve the genome-wide significance. These undiscovered variants further limit the prediction capability of GWAS. Restricted access to the individual-level data and the increasing availability of the published GWAS results motivate the development of methods integrating both the individual-level and summary-level data. How to build the connection between the individual-level and summary-level data determines the efficiency of using the existing abundant summary-level resources with limited individual-level data, and this issue inspires more efforts in the existing area. RESULTS: In this study, we propose a novel statistical approach, LEP, which provides a novel way of modeling the connection between the individual-level data and summary-level data. LEP integrates both types of data by LEveraging Pleiotropy to increase the statistical power of risk variants identification and the accuracy of risk prediction. The algorithm for parameter estimation is developed to handle genome-wide-scale data. Through comprehensive simulation studies, we demonstrated the advantages of LEP over the existing methods. We further applied LEP to perform integrative analysis of Crohn's disease from WTCCC and summary statistics from GWAS of some other diseases, such as Type 1 diabetes, Ulcerative colitis and Primary biliary cirrhosis. LEP was able to significantly increase the statistical power of identifying risk variants and improve the risk prediction accuracy from 63.39% (±0.58%) to 68.33% (±0.32%) using about 195 000 variants. AVAILABILITY AND IMPLEMENTATION: The LEP software is available at https://github.com/daviddaigithub/LEP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mingwei Dai, Yao Wang 0003, Yue Liu 0035, Jin Liu 0011, Zongben Xu, Can Yang 0002
Bioinform.7
2019 A Graph-Based Semisupervised Deep Learning Model for PolSAR Image Classification
abstract
Aiming at improving the classification accuracy with limited numbers of labeled pixels in polarimetric synthetic aperture radar (PolSAR) image classification task, this paper presents a graph-based semisupervised deep learning model for PolSAR image classification. It models the PolSAR image as an undirected graph, where the nodes correspond to the labeled and unlabeled pixels, and the weighted edges represent similarities between the pixels. Upon the graph, we design an energy function incorporating a semisupervision term, a convolutional neural network (CNN) term, and a pairwise smoothness term. The employed CNN extracts abstract and data-driven polarimetric features and outputs class label predictions to the graph model. The semisupervision term enforces the category label constraints on the human-labeled pixels. The pairwise smoothness term encourages class label smoothness and the alignment of class label boundaries with the image edges. Starting from an initialized class label map generated based on K-Wishart distribution hypothesis or superpixel segmentation of PauliRGB images, we iteratively and alternately optimize the defined energy function until it converges. We conducted experiments on two real benchmark PolSAR images, and extensive experiments demonstrated that our approach achieved the state-of-the-art results for PolSAR image classification.
Haixia Bi, Jian Sun 0009, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.3
2019 An Active Deep Learning Approach for Minimally Supervised PolSAR Image Classification
abstract
Recently, deep neural networks have received intense interests in polarimetric synthetic aperture radar (PolSAR) image classification. However, its success is subject to the availability of large amounts of annotated data which require great efforts of experienced human annotators. Aiming at improving the classification performance with greatly reduced annotation cost, this paper presents an active deep learning approach for minimally supervised PolSAR image classification, which integrates active learning and fine-tuned convolutional neural network (CNN) into a principled framework. Starting from a CNN trained using a very limited number of labeled pixels, we iteratively and actively select the most informative candidates for annotation, and incrementally fine-tune the CNN by incorporating the newly annotated pixels. Moreover, to boost the performance and robustness of the proposed method, we employ Markov random field (MRF) to enforce class label smoothness, and data augmentation technique to enlarge the training set. We conducted extensive experiments on four real benchmark PolSAR images, and experiments demonstrated that our approach achieved state-of-the-art classification results with significantly reduced annotation cost.
Haixia Bi, Feng Xu 0001, Zhiqiang Wei 0004, Yong Xue, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.5
2019 An Optimized Deep Network Representation of Multimutation Differential Evolution and its Application in Seismic Inversion
abstract
Seismic inversion problems are well-known to be nonlinear and their misfit functions often involve many local minima. Global optimization methods are capable of converging to the global minimum of a misfit function, thus, they are promising in seismic inversion. As a global optimization method, multimutation differential evolution (MMDE) has been proven to be effective in solving high-dimensional seismic inversion problems. However, it is challenging to choose the optimal parameters for MMDE to achieve the best performance in seismic inversion. In this paper, we propose a new deep network based on MMDE and name it as MMDE-Net, which enables us to learn the optimal parameters by using a network training procedure rather than empirically choosing them. Benefiting from the learned parameters, MMDE-Net has advantages over MMDE in applications. Numerical examples based on synthetic and field data set clearly indicate that MMDE-Net can provide faster convergence speed and better inversion result than conventional methods in seismic inversion.
Zhaoqi Gao, Zhibin Pan, Jinghuai Gao, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.5
2019 Nonconvex-Sparsity and Nonlocal-Smoothness-Based Blind Hyperspectral Unmixing
abstract
Blind hyperspectral unmixing (HU), as a crucial technique for hyperspectral data exploitation, aims to decompose mixed pixels into a collection of constituent materials weighted by the corresponding fractional abundances. In recent years, nonnegative matrix factorization (NMF) based methods have become more and more popular for this task and achieved promising performance. Among these methods, two types of properties upon the abundances, namely the sparseness and the structural smoothness, have been explored and shown to be important for blind HU. However, all of previous methods ignores another important insightful property possessed by a natural hyperspectral images (HSI), non-local smoothness, which means that similar patches in a larger region of an HSI are sharing the similar smoothness structure. Based on previous attempts on other tasks, such a prior structure reflects intrinsic configurations underlying a HSI, and is thus expected to largely improve the performance of the investigated HU problem. In this paper, we firstly consider such prior in HSI by encoding it as the nonlocal total variation (NLTV) regularizer. Furthermore, by fully exploring the intrinsic structure of HSI, we generalize NLTV to non-local HSI TV (NLHTV) to make the model more suitable for the bind HU task. By incorporating these two regularizers, together with a non-convex log-sum form regularizer characterizing the sparseness of abundance maps, to the NMF model, we propose novel blind HU models named NLTV/NLHTV and log-sum regularized NMF (NLTV-LSRNMF/NLHTV-LSRNMF), respectively. To solve the proposed models, an efficient algorithm is designed based on alternative optimization strategy (AOS) and alternating direction method of multipliers (ADMM). Extensive experiments conducted on both simulated and real hyperspectral data sets substantiate the superiority of the proposed approach over other competing ones for blind HU task.
Jing Yao 0002, Deyu Meng, Qian Zhao 0002, Wenfei Cao, Zongben Xu
IEEE Trans. Image Process.5
2019 Optimizing a Parameterized Plug-and-Play ADMM for Iterative Low-Dose CT Reconstruction
abstract
Reducing the exposure to X-ray radiation while maintaining a clinically acceptable image quality is desirable in various CT applications. To realize low-dose CT (LdCT) imaging, model-based iterative reconstruction (MBIR) algorithms are widely adopted, but they require proper prior knowledge assumptions in the sinogram and/or image domains and involve tedious manual optimization of multiple parameters. In this paper, we propose a deep learning (DL)-based strategy for MBIR to simultaneously address prior knowledge design and MBIR parameter selection in one optimization framework. Specifically, a parameterized plug-and-play alternating direction method of multipliers (3pADMM) is proposed for the general penalized weighted least-squares model, and then, by adopting the basic idea of DL, the parameterized plug-and-play (3p) prior and the related parameters are optimized simultaneously in a single framework using a large number of training data. The main contribution of this paper is that the 3p prior and the related parameters in the proposed 3pADMM framework can be supervised and optimized simultaneously to achieve robust LdCT reconstruction performance. Experimental results obtained on clinical patient datasets demonstrate that the proposed method can achieve promising gains over existing algorithms for LdCT image reconstruction in terms of noise-induced artifact suppression and edge detail preservation.
Ji He 0001, Yan Yang 0007, Dong Zeng, Zhaoying Bian, Hao Zhang 0026, Jian Sun 0009, Zongben Xu, Jianhua Ma 0001
IEEE Trans. Medical Imaging8
2019 An Efficient Iterative Cerebral Perfusion CT Reconstruction via Low-Rank Tensor Decomposition With Spatial-Temporal Total Variation Regularization
abstract
Cerebrovascular diseases, i.e., acute stroke, are a common cause of serious long-term disability. Cerebral perfusion computed tomography (CPCT) can provide rapid, high-resolution, quantitative hemodynamic maps to assess and stratify perfusion in patients with acute stroke symptoms. However, CPCT imaging typically involves a substantial radiation dose due to its repeated scanning protocol. Therefore, in this paper, we present a low-dose CPCT image reconstruction method to yield high-quality CPCT images and high-precision hemodynamic maps by utilizing the great similarity information among the repeated scanned CPCT images. Specifically, a newly developed low-rank tensor decomposition with spatial-temporal total variation (LRTD-STTV) regularization is incorporated into the reconstruction model. In the LRTD-STTV regularization, the tensor Tucker decomposition is used to describe global spatial-temporal correlations hidden in the sequential CPCT images, and it is superior to the matricization model (i.e., low-rank model) that fails to fully investigate the prior knowledge of the intrinsic structures of the CPCT images after vectorizing the CPCT images. Moreover, the spatial-temporal TV regularization is used to characterize the local piecewise smooth structure in the spatial domain and the pixels' similarity with the adjacent frames in the temporal domain, because the intensity at each pixel in CPCT images is similar to its neighbors. Therefore, the presented LRTD-STTV model can efficiently deliver faithful underlying information of the CPCT images and preserve the spatial structures. An efficient alternating direction method of multipliers algorithm is also developed to solve the presented LRTD-STTV model. Extensive experimental results on numerical phantom and patient data are clearly demonstrated that the presented model can significantly improve the quality of CPCT images and provide accurate diagnostic features in hemodynamic maps for low-dose cases compared with the existing popular algorithms.
Sui Li, Dong Zeng, Jiangjun Peng, Zhaoying Bian, Hao Zhang 0026, Qi Xie 0002, Yuting Liao, Shanli Zhang, Jing Huang 0018, Deyu Meng, Zongben Xu, Jianhua Ma 0001
IEEE Trans. Medical Imaging12
2018 Margin Based PU Learning
abstract
The PU learning problem concerns about learning from positive and unlabeled data. A popular heuristic is to iteratively enlarge training set based on some margin-based criterion. However, little theoretical analysis has been conducted to support the success of these heuristic methods. In this work, we show that not all margin-based heuristic rules are able to improve the learned classifiers iteratively. We find that a so-called large positive margin oracle is necessary to guarantee the success of PU learning. Under this oracle, a provable positive-margin based PU learning algorithm is proposed for linear regression and classification under the truncated Gaussian distributions. The proposed algorithm is able to reduce the recovering error geometrically proportional to the positive margin. Extensive experiments on real-world datasets verify our theory and the state-of-the-art performance of the proposed PU learning algorithm.
Tieliang Gong, Guangtao Wang, Jieping Ye, Zongben Xu
AAAI4
2018 Convergence of multi-block Bregman ADMM for nonconvex composite problems
Fenghui Wang, Wenfei Cao, Zongben Xu
Sci. China Inf. Sci.3
2018 Diverse lesion detection from retinal images by subspace learning over normal samples
Benzhi Chen, Lisheng Wang, Jian Sun 0009, Huai Chen, Yinghua Fu, Shouren Lan, Zongben Xu
Neurocomputing8
2018 Robust subspace clustering via penalized mixture of Gaussians
Jing Yao 0002, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu
Neurocomputing5
2018 Neural multi-atlas label fusion: Application to cardiac MR images
Heran Yang, Jian Sun 0009, Huibin Li 0001, Lisheng Wang, Zongben Xu
Medical Image Anal.5
2018 Kronecker-Basis-Representation Based Tensor Sparsity and Its Applications to Tensor Recovery
abstract
As a promising way for analyzing data, sparse modeling has achieved great success throughout science and engineering. It is well known that the sparsity/low-rank of a vector/matrix can be rationally measured by nonzero-entries-number ( norm)/nonzero- singular-values-number (rank), respectively. However, data from real applications are often generated by the interaction of multiple factors, which obviously cannot be sufficiently represented by a vector/matrix, while a high order tensor is expected to provide more faithful representation to deliver the intrinsic structure underlying such data ensembles. Unlike the vector/matrix case, constructing a rational high order sparsity measure for tensor is a relatively harder task. To this aim, in this paper we propose a measure for tensor sparsity, called Kronecker-basis-representation based tensor sparsity measure (KBR briefly), which encodes both sparsity insights delivered by Tucker and CANDECOMP/PARAFAC (CP) low-rank decompositions for a general tensor. Then we study the KBR regularization minimization (KBRM) problem, and design an effective ADMM algorithm for solving it, where each involved parameter can be updated with closed-form equations. Such an efficient solver makes it possible to extend KBR to various tasks like tensor completion and tensor robust principal component analysis. A series of experiments, including multispectral image (MSI) denoising, MSI completion and background subtraction, substantiate the superiority of the proposed methods beyond state-of-the-arts.
Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Denoising Hyperspectral Image With Non-i.i.d. Noise Structure
abstract
Hyperspectral image (HSI) denoising has been attracting much research attention in remote sensing area due to its importance in improving the HSI qualities. The existing HSI denoising methods mainly focus on specific spectral and spatial prior knowledge in HSIs, and share a common underlying assumption that the embedded noise in HSI is independent and identically distributed (i.i.d.). In real scenarios, however, the noise existed in a natural HSI is always with much more complicated non-i.i.d. statistical structures and the under-estimation to this noise complexity often tends to evidently degenerate the robustness of current methods. To alleviate this issue, this paper attempts the first effort to model the HSI noise using a non-i.i.d. mixture of Gaussians (NMoGs) noise assumption, which finely accords with the noise characteristics possessed by a natural HSI and thus is capable of adapting various practical noise shapes. Then we integrate such noise modeling strategy into the low-rank matrix factorization (LRMF) model and propose an NMoG-LRMF model in the Bayesian framework. A variational Bayes algorithm is then designed to infer the posterior of the proposed model. As substantiated by our experiments implemented on synthetic and real noisy HSIs, the proposed method performs more robust beyond the state-of-the-arts.
Yang Chen 0057, Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Zongben Xu
IEEE Trans. Cybern.5
2018 Greedy Criterion in Orthogonal Greedy Learning
abstract
Orthogonal greedy learning (OGL) is a stepwise learning scheme that starts with selecting a new atom from a specified dictionary via the steepest gradient descent (SGD) and then builds the estimator through orthogonal projection. In this paper, we found that SGD is not the unique greedy criterion and introduced a new greedy criterion, called as " -greedy threshold" for learning. Based on this new greedy criterion, we derived a straightforward termination rule for OGL. Our theoretical study shows that the new learning scheme can achieve the existing (almost) optimal learning rate of OGL. Numerical experiments are also provided to support that this new scheme can achieve almost optimal generalization performance while requiring less computation than OGL.
Lin Xu 0001, Shaobo Lin, Jinshan Zeng, Yi Fang 0006, Zongben Xu
IEEE Trans. Cybern.6
2018 Hyperspectral Image Classification With Markov Random Fields and a Convolutional Neural Network
abstract
This paper presents a new supervised classification algorithm for remotely sensed hyperspectral image (HSI) which integrates spectral and spatial information in a unified Bayesian framework. First, we formulate the HSI classification problem from a Bayesian perspective. Then, we adopt a convolutional neural network (CNN) to learn the posterior class distributions using a patch-wise training strategy to better use the spatial information. Next, spatial information is further considered by placing a spatial smoothness prior on the labels. Finally, we iteratively update the CNN parameters using stochastic gradient decent and update the class labels of all pixel vectors using -expansion min-cut-based algorithm. Compared with the other state-of-the-art methods, the classification method achieves better performance on one synthetic data set and two benchmark HSI data sets in a number of experimental settings.
Xiangyong Cao, Feng Zhou 0001, Lin Xu 0001, Deyu Meng, Zongben Xu, John W. Paisley
IEEE Trans. Image Process.5
2018 A Preference-Based Multiobjective Evolutionary Approach for Sparse Optimization
abstract
Iterative thresholding is a dominating strategy for sparse optimization problems. The main goal of iterative thresholding methods is to find a so-called -sparse solution. However, the setting of regularization parameters or the estimation of the true sparsity are nontrivial in iterative thresholding methods. To overcome this shortcoming, we propose a preference-based multiobjective evolutionary approach to solve sparse optimization problems in compressive sensing. Our basic strategy is to search the knee part of weakly Pareto front with preference on the true -sparse solution. In the noiseless case, it is easy to locate the exact position of the -sparse solution from the distribution of the solutions found by our proposed method. Therefore, our method has the ability to detect the true sparsity. Moreover, any iterative thresholding methods can be used as a local optimizer in our proposed method, and no prior estimation of sparsity is required. The proposed method can also be extended to solve sparse optimization problems with noise. Extensive experiments have been conducted to study its performance on artificial signals and magnetic resonance imaging signals. Our experimental results have shown that our proposed method is very effective for detecting sparsity and can improve the reconstruction ability of existing iterative thresholding methods.
Hui Li 0020, Qingfu Zhang 0001, Jingda Deng, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2017 Should We Encode Rain Streaks in Video as Deterministic or Stochastic?
abstract
Videos taken in the wild sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal in a video (RSRV) is thus an important issue and has been attracting much attention in computer vision. Different from previous RSRV methods formulating rain streaks as a deterministic message, this work first encodes the rains in a stochastic manner, i.e., a patch-based mixture of Gaussians. Such modification makes the proposed model capable of finely adapting a wider range of rain variations instead of certain types of rain configurations as traditional. By integrating with the spatiotemporal smoothness configuration of moving objects and low-rank structure of background scene, we propose a concise model for RSRV, containing one likelihood term imposed on the rain streak layer and two prior terms on the moving object and background scene layers of the video. Experiments implemented on videos with synthetic and real rains verify the superiority of the proposed method, as compared with the state-of-the-art methods, both visually and quantitatively in various performance metrics.
Wei Wei 0006, Lixuan Yi, Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu
ICCV6
2017 Polsar image classification based on three-dimensional wavelet texture features and Markov random field
abstract
The speckle effect embedded in polarimetric synthetic aperture radar (PolSAR) data damages the performance of PolSAR image classification greatly. To alleviate this issue, a new supervised classification method, which introduces spatial consistency in both feature extraction and classification steps is proposed. Specifically, three-dimensional discrete wavelet transform (3D-DWT) is used to extract spectral-spatial texture features, which are proved to be more discriminative than original ones. Afterward, label smoothness prior is incorporated in the classification, which is implemented using a Markov random field (MRF). To demonstrate the validity of the proposed method, real PolSAR image is used in experiments. Compared with the other state-of-the-art methods, this method achieves higher classification accuracy and better visual spatial connectivity.
Haixia Bi, Lin Xu 0001, Xiangyong Cao, Zongben Xu
IGARSS4
2017 IGESS: a statistical approach to integrating individual-level genotype data and summary statistics in genome-wide association studies
abstract
MOTIVATION: Results from genome-wide association studies (GWAS) suggest that a complex phenotype is often affected by many variants with small effects, known as 'polygenicity'. Tens of thousands of samples are often required to ensure statistical power of identifying these variants with small effects. However, it is often the case that a research group can only get approval for the access to individual-level genotype data with a limited sample size (e.g. a few hundreds or thousands). Meanwhile, summary statistics generated using single-variant-based analysis are becoming publicly available. The sample sizes associated with the summary statistics datasets are usually quite large. How to make the most efficient use of existing abundant data resources largely remains an open question. RESULTS: In this study, we propose a statistical approach, IGESS, to increasing statistical power of identifying risk variants and improving accuracy of risk prediction by i ntegrating individual level ge notype data and s ummary s tatistics. An efficient algorithm based on variational inference is developed to handle the genome-wide analysis. Through comprehensive simulation studies, we demonstrated the advantages of IGESS over the methods which take either individual-level data or summary statistics data as input. We applied IGESS to perform integrative analysis of Crohns Disease from WTCCC and summary statistics from other studies. IGESS was able to significantly increase the statistical power of identifying risk variants and improve the risk prediction accuracy from 63.2% ( ±0.4% ) to 69.4% ( ±0.1% ) using about 240 000 variants. AVAILABILITY AND IMPLEMENTATION: The IGESS software is available at https://github.com/daviddaigithub/IGESS . CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mingwei Dai, Jingsi Ming, Mingxuan Cai, Jin Liu 0011, Can Yang 0002, Zongben Xu
Bioinform.7
2017 Integration of 3-dimensional discrete wavelet transform and Markov random field for hyperspectral image classification
Xiangyong Cao, Lin Xu 0001, Deyu Meng, Qian Zhao 0002, Zongben Xu
Neurocomputing5
2017 Generalization Analysis of Fredholm Kernel Regularized Classifiers
abstract
Recently, a new framework, Fredholm learning, was proposed for semisupervised learning problems based on solving a regularized Fredholm integral equation. It allows a natural way to incorporate unlabeled data into learning algorithms to improve their prediction performance. Despite rapid progress on implementable algorithms with theoretical guarantees, the generalization ability of Fredholm kernel learning has not been studied. In this letter, we focus on investigating the generalization performance of a family of classification algorithms, referred to as Fredholm kernel regularized classifiers. We prove that the corresponding learning rate can achieve [Formula: see text] ([Formula: see text] is the number of labeled samples) in a limiting case. In addition, a representer theorem is provided for the proposed regularized scheme, which underlies its applications.
Tieliang Gong, Zongben Xu, Hong Chen 0004
Neural Comput.2
2017 Unsupervised PolSAR Image Classification Using Discriminative Clustering
abstract
This paper presents a novel unsupervised image classification method for polarimetric synthetic aperture radar (PolSAR) data. The proposed method is based on a discriminative clustering framework that explicitly relies on a discriminative supervised classification technique to perform unsupervised clustering. To implement this idea, we design an energy function for unsupervised PolSAR image classification by combining a supervised softmax regression model with a Markov random field smoothness constraint. In this model, both the pixelwise class labels and classifiers are taken as unknown variables to be optimized. Starting from the initialized class labels generated by Cloude-Pottier decomposition and $K$ -Wishart distribution hypothesis, we iteratively optimize the classifiers and class labels by alternately minimizing the energy function with respect to them. Finally, the optimized class labels are taken as the classification result, and the classifiers for different classes are also derived as a side effect. We apply this approach to real PolSAR benchmark data. Extensive experiments justify that our approach can effectively classify the PolSAR image in an unsupervised way and produce higher accuracies than the compared state-of-the-art methods.
Haixia Bi, Jian Sun 0009, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.3
2017 Robust Low-Dose CT Sinogram Preprocessing via Exploiting Noise-Generating Mechanism
abstract
Computed tomography (CT) image recovery from low-mAs acquisitions without adequate treatment is always severely degraded due to a number of physical factors. In this paper, we formulate the low-dose CT sinogram preprocessing as a standard maximum a posteriori (MAP) estimation, which takes full consideration of the statistical properties of the two intrinsic noise sources in low-dose CT, i.e., the X-ray photon statistics and the electronic noise background. In addition, instead of using a general image prior as found in the traditional sinogram recovery models, we design a new prior formulation to more rationally encode the piecewise-linear configurations underlying a sinogram than previously used ones, like the TV prior term. As compared with the previous methods, especially the MAP-based ones, both the likelihood/loss and prior/regularization terms in the proposed model are ameliorated in a more accurate manner and better comply with the statistical essence of the generation mechanism of a practical sinogram. We further construct an efficient alternating direction method of multipliers algorithm to solve the proposed MAP framework. Experiments on simulated and real low-dose CT data demonstrate the superiority of the proposed method according to both visual inspection and comprehensive quantitative performance evaluation.
Qi Xie 0002, Dong Zeng, Qian Zhao 0002, Deyu Meng, Zongben Xu, Zhengrong Liang, Jianhua Ma 0001
IEEE Trans. Medical Imaging5
2017 Low-Dose Dynamic Cerebral Perfusion Computed Tomography Reconstruction via Kronecker-Basis-Representation Tensor Sparsity Regularization
abstract
Dynamic cerebral perfusion computed tomography (DCPCT) has the ability to evaluate the hemodynamic information throughout the brain. However, due to multiple 3-D image volume acquisitions protocol, DCPCT scanning imposes high radiation dose on the patients with growing concerns. To address this issue, in this paper, based on the robust principal component analysis (RPCA, or equivalently the low-rank and sparsity decomposition) model and the DCPCT imaging procedure, we propose a new DCPCT image reconstruction algorithm to improve low-dose DCPCT and perfusion maps quality via using a powerful measure, called Kronecker-basis-representation tensor sparsity regularization, for measuring low-rankness extent of a tensor. For simplicity, the first proposed model is termed tensor-based RPCA (T-RPCA). Specifically, the T-RPCA model views the DCPCT sequential images as a mixture of low-rank, sparse, and noise components to describe the maximum temporal coherence of spatial structure among phases in a tensor framework intrinsically. Moreover, the low-rank component corresponds to the "background" part with spatial-temporal correlations, e.g., static anatomical contribution, which is stationary over time about structure, and the sparse component represents the time-varying component with spatial-temporal continuity, e.g., dynamic perfusion enhanced information, which is approximately sparse over time. Furthermore, an improved nonlocal patch-based T-RPCA (NL-T-RPCA) model which describes the 3-D block groups of the "background" in a tensor is also proposed. The NL-T-RPCA model utilizes the intrinsic characteristics underlying the DCPCT images, i.e., nonlocal self-similarity and global correlation. Two efficient algorithms using alternating direction method of multipliers are developed to solve the proposed T-RPCA and NL-T-RPCA models, respectively. Extensive experiments with a digital brain perfusion phantom, preclinical monkey data, and clinical patient data clearly demonstrate that the two proposed models can achieve more gains than the existing popular algorithms in terms of both quantitative and visual quality evaluations from low-dose acquisitions, especially as low as 20 mAs.
Dong Zeng, Qi Xie 0002, Wenfei Cao, Jiahui Lin, Hao Zhang 0026, Shanli Zhang, Jing Huang 0018, Zhaoying Bian, Deyu Meng, Zongben Xu, Zhengrong Liang, Wufan Chen, Jianhua Ma 0001
IEEE Trans. Medical Imaging10
2017 Multimodal 2D+3D Facial Expression Recognition With Deep Fusion Convolutional Neural Network
abstract
This paper presents a novel and efficient deep fusion convolutional neural network (DF-CNN) for multimodal 2D+3D facial expression recognition (FER). DF-CNN comprises a feature extraction subnet, a feature fusion subnet, and a softmax layer. In particular, each textured three-dimensional (3D) face scan is represented as six types of 2D facial attribute maps (i.e., geometry map, three normal maps, curvature map, and texture map), all of which are jointly fed into DF-CNN for feature learning and fusion learning, resulting in a highly concentrated facial representation (32-dimensional). Expression prediction is performed by two ways: 1) learning linear support vector machine classifiers using the 32-dimensional fused deep features, or 2) directly performing softmax prediction using the six-dimensional expression probability vectors. Different from existing 3D FER methods, DF-CNN combines feature learning and fusion learning into a single end-to-end training framework. To demonstrate the effectiveness of DF-CNN, we conducted comprehensive experiments to compare the performance of DFCNN with handcrafted features, pre-trained deep features, finetuned deep features, and state-of-the-art methods on three 3D face datasets (i.e., BU-3DFE Subset I, BU-3DFE Subset II, and Bosphorus Subset). In all cases, DF-CNN consistently achieved the best results. To the best of our knowledge, this is the first work of introducing deep CNN to 3D FER and deep learning-based featurelevel fusion for multimodal 2D+3D FER.
Huibin Li 0001, Jian Sun 0009, Zongben Xu, Liming Chen 0002
IEEE Trans. Multim.3
2017 Graph PCA Hashing for Similarity Search
abstract
This paper proposes a new hashing framework to conduct similarity search via the following steps: first, employing linear clustering methods to obtain a set of representative data points and a set of landmarks of the big dataset; second, using the landmarks to generate a probability representation for each data point. The proposed probability representation method is further proved to preserve the neighborhood of each data point. Third, PCA is integrated with manifold learning to lean the hash functions using the probability representations of all representative data points. As a consequence, the proposed hashing method achieves efficient similarity search (with linear time complexity) and effective hashing performance and high generalization ability (simultaneously preserving two kinds of complementary similarity structures, i.e., local structures via manifold learning and global structures via PCA). Experimental results on four public datasets clearly demonstrate the advantages of our proposed method in terms of similarity search, compared to the state-of-the-art hashing methods.
Xiaofeng Zhu 0001, Xuelong Li 0001, Shichao Zhang 0001, Zongben Xu, Litao Yu, Can Wang 0004
IEEE Trans. Multim.4
2017 Shrinkage Degree in L2-Rescale Boosting for Regression
abstract
L2-rescale boosting (L2-RBoosting) is a variant of L2-Boosting, which can essentially improve the generalization performance of L2-Boosting. The key feature of L2-RBoosting lies in introducing a shrinkage degree to rescale the ensemble estimate in each iteration. Thus, the shrinkage degree determines the performance of L2-RBoosting. The aim of this paper is to develop a concrete analysis concerning how to determine the shrinkage degree in L2-RBoosting. We propose two feasible ways to select the shrinkage degree. The first one is to parameterize the shrinkage degree and the other one is to develop a data-driven approach. After rigorously analyzing the importance of the shrinkage degree in L2-RBoosting, we compare the pros and cons of the proposed methods. We find that although these approaches can reach the same learning rates, the structure of the final estimator of the parameterized approach is better, which sometimes yields a better generalization capability when the number of sample is finite. With this, we recommend to parameterize the shrinkage degree of L2-RBoosting. We also present an adaptive parameter-selection strategy for shrinkage degree and verify its feasibility through both theoretical analysis and numerical verification. The obtained results enhance the understanding of L2-RBoosting and give guidance on how to use it for regression tasks.
Lin Xu 0001, Shaobo Lin, Yao Wang 0003, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2017 Improving Separability of Structures with Similar Attributes in 2D Transfer Function Design
abstract
The 2D transfer function based on scalar value and gradient magnitude (SG-TF) is popularly used in volume rendering. However, it is plagued by the boundary-overlapping problem: different structures with similar attributes have the same region in SG-TF space, and their boundaries are usually connected. The SG-TF thus often fails in separating these structures (or their boundaries) and has limited ability to classify different objects in real-world 3D images. To overcome such a difficulty, we propose a novel method for boundary separation by integrating spatial connectivity computation of the boundaries and set operations on boundary voxels into the SG-TF. Specifically, spatial positions of boundaries and their regions in the SG-TF space are computed, from which boundaries can be well separated and volume rendered in different colors. In the method, the boundaries are divided into three classes and different boundary-separation techniques are applied to them, respectively. The complex task of separating various boundaries in 3D images is then simplified by breaking it into several small separation problems. The method shows good object classification ability in real-world 3D images while avoiding the complexity of high-dimensional transfer functions. Its effectiveness and validation is demonstrated by many experimental results to visualize boundaries of different structures in complex real-world 3D images.
Shouren Lan, Lisheng Wang, Yipeng Song, Yu-Ping Wang 0002, Liping Yao, Zongben Xu
IEEE Trans. Vis. Comput. Graph.8
2016 A multi-phase multiobjective approach based on decomposition for sparse reconstruction
abstract
Solving sparse optimization problems via regularization frameworks is the dominant methodology for reconstructing sparse signals in the area of compressive sensing. In recent a few years, the use of multiobjective evolutionary algorithms (MOEAs) for sparse optimization has also attracted some research interests. Under the multiobjective framework, the loss term (error) and the regularization term (sparsity) are treated as two separate objective functions. So far, two popular multiobjective frameworks, NSGA-II and MOEA/D, have been used for sparse optimization. In this paper, we further develop a new MOEA/D variant for sparse reconstruction and sparsity detection, which involves three phases - approximating Pareto front (PF) in a chain order (phase 1) and in a random order (phase 2), and exploiting a knee region (phase 3 - optional). Our experimental results show that our proposed method is more effective than the earlier version of MOEA/D and the HALF solver in sparse signal reconstruction and sparsity detection.
Hui Li 0020, Qingfu Zhang 0001, Zongben Xu, Jingda Deng
CEC4
2016 Multispectral Images Denoising by Intrinsic Tensor Sparsity Regularization
abstract
Multispectral images (MSI) can help deliver more faithful representation for real scenes than the traditional image system, and enhance the performance of many computer vision tasks. In real cases, however, an MSI is always corrupted by various noises. In this paper, we propose a new tensor-based denoising approach by fully considering two intrinsic characteristics underlying an MSI, i.e., the global correlation along spectrum (GCS) and nonlocal self-similarity across space (NSS). In specific, we construct a new tensor sparsity measure, called intrinsic tensor sparsity (ITS) measure, which encodes both sparsity insights delivered by the most typical Tucker and CANDECOMP/ PARAFAC (CP) low-rank decomposition for a general tensor. Then we build a new MSI denoising model by applying the proposed ITS measure on tensors formed by non-local similar patches within the MSI. The intrinsic GCS and NSS knowledge can then be efficiently explored under the regularization of this tensor sparsity measure to finely rectify the recovery of a MSI from its corruption. A series of experiments on simulated and real MSI denoising problems show that our method outperforms all state-of-the-arts under comprehensive quantitative performance measures.
Qi Xie 0002, Qian Zhao 0002, Deyu Meng, Zongben Xu, Shuhang Gu, Wangmeng Zuo, Lei Zhang 0006
CVPR4
2016 Deep Fusion Net for Multi-atlas Segmentation: Application to Cardiac MR Images
Heran Yang, Jian Sun 0009, Huibin Li 0001, Lisheng Wang, Zongben Xu
MICCAI (2)5
2016 Deep ADMM-Net for Compressive Sensing MRI
abstract
Compressive Sensing (CS) is an effective approach for fast Magnetic Resonance Imaging (MRI). It aims at reconstructing MR image from a small number of under-sampled data in k-space, and accelerating the data acquisition in MRI. To improve the current MRI system in reconstruction accuracy and computational speed, in this paper, we propose a novel deep architecture, dubbed ADMM-Net. ADMM-Net is defined over a data flow graph, which is derived from the iterative procedures in Alternating Direction Method of Multipliers (ADMM) algorithm for optimizing a CS-based MRI model. In the training phase, all parameters of the net, e.g., image transforms, shrinkage functions, etc., are discriminatively trained end-to-end using L-BFGS algorithm. In the testing phase, it has computational overhead similar to ADMM but uses optimized parameters learned from the training data for CS-based reconstruction task. Experiments on MRI image reconstruction under different sampling ratios in k-space demonstrate that it significantly improves the baseline ADMM algorithm and achieves high reconstruction accuracies with fast computational speed.
Yan Yang 0007, Jian Sun 0009, Huibin Li 0001, Zongben Xu
NIPS4
2016 Joint sparse canonical correlation analysis for detecting differential imaging genetics modules
abstract
MOTIVATION: Imaging genetics combines brain imaging and genetic information to identify the relationships between genetic variants and brain activities. When the data samples belong to different classes (e.g. disease status), the relationships may exhibit class-specific patterns that can be used to facilitate the understanding of a disease. Conventional approaches often perform separate analysis on each class and report the differences, but ignore important shared patterns. RESULTS: In this paper, we develop a multivariate method to analyze the differential dependency across multiple classes. We propose a joint sparse canonical correlation analysis method, which uses a generalized fused lasso penalty to jointly estimate multiple pairs of canonical vectors with both shared and class-specific patterns. Using a data fusion approach, the method is able to detect differentially correlated modules effectively and efficiently. The results from simulation studies demonstrate its higher accuracy in discovering both common and differential canonical correlations compared to conventional sparse CCA. Using a schizophrenia dataset with 92 cases and 116 controls including a single nucleotide polymorphism (SNP) array and functional magnetic resonance imaging data, the proposed method reveals a set of distinct SNP-voxel interaction modules for the schizophrenia patients, which are verified to be both statistically and biologically significant. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at https://sites.google.com/site/jianfang86/JSCCA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Jian Fang 0001, Dongdong Lin, S. Charles Schulz, Zongben Xu, Vince D. Calhoun, Yu-Ping Wang 0002
Bioinform.4
2016 Learning capability of the truncated greedy algorithm
Lin Xu 0001, Shaobo Lin, Zongben Xu
Sci. China Inf. Sci.3
2016 Block-sparse compressed sensing with partially known signal support via non-convex minimisation
abstract
The mixed l 2 / l p (0 < p ≤ 1) norm minimisation method with partially known support for recovering block‐sparse signals is studied. The authors mainly extend this work on block‐sparse compressed sensing by incorporating some known part of the block support information as a priori and establish sufficient restricted p ‐isometry property ( p ‐RIP) conditions for exact and robust recovery. The authors’ theoretical results show it is possible to recover the block‐sparse signals via l 2 / l p minimisation from reduced number of measurements by applying the partially known support. The authors also derive a lower bound on necessary random Gaussian measurements for the p ‐RIP conditions to hold with high possibility. Finally, a series of numerical experiments are carried out to illustrate that fewer measurements with smaller p are needed to reconstruct the signal.
Shiying He, Yao Wang 0003, Jianjun Wang 0003, Zongben Xu
IET Signal Process.4
2016 Learning and approximation capabilities of orthogonal super greedy algorithm
Jian Fang 0001, Shaobo Lin, Zongben Xu
Knowl. Based Syst.3
2016 Re-scale AdaBoost for attack detection in collaborative filtering recommender systems
Zhihai Yang, Lin Xu 0001, Zhongmin Cai, Zongben Xu
Knowl. Based Syst.4
2016 Learning With ℓ1-Regularizer Based on Markov Resampling
abstract
Learning with l1 -regularizer has brought about a great deal of research in learning theory community. Previous known results for the learning with l1 -regularizer are based on the assumption that samples are independent and identically distributed (i.i.d.), and the best obtained learning rate for the l1 -regularization type algorithms is O(1/√m) , where m is the samples size. This paper goes beyond the classic i.i.d. framework and investigates the generalization performance of least square regression with l1 -regularizer ( l1 -LSR) based on uniformly ergodic Markov chain (u.e.M.c) samples. On the theoretical side, we prove that the learning rate of l1 -LSR for u.e.M.c samples l1 -LSR(M) is with the order of O(1/m) , which is faster than O(1/√m) for the i.i.d. counterpart. On the practical side, we propose an algorithm based on resampling scheme to generate u.e.M.c samples. We show that the proposed l1 -LSR(M) improves on the l1 -LSR(i.i.d.) in generalization error at the low cost of u.e.M.c resampling.
Tieliang Gong, Bin Zou 0002, Zongben Xu
IEEE Trans. Cybern.3
2016 Total Variation Regularized Tensor RPCA for Background Subtraction From Compressive Measurements
abstract
Background subtraction has been a fundamental and widely studied task in video analysis, with a wide range of applications in video surveillance, teleconferencing, and 3D modeling. Recently, motivated by compressive imaging, background subtraction from compressive measurements (BSCM) is becoming an active research task in video surveillance. In this paper, we propose a novel tensor-based robust principal component analysis (TenRPCA) approach for BSCM by decomposing video frames into backgrounds with spatial-temporal correlations and foregrounds with spatio-temporal continuity in a tensor framework. In this approach, we use 3D total variation to enhance the spatio-temporal continuity of foregrounds, and Tucker decomposition to model the spatio-temporal correlations of video background. Based on this idea, we design a basic tensor RPCA model over the video frames, dubbed as the holistic TenRPCA model. To characterize the correlations among the groups of similar 3D patches of video background, we further design a patch-group-based tensor RPCA model by joint tensor Tucker decompositions of 3D patch groups for modeling the video background. Efficient algorithms using the alternating direction method of multipliers are developed to solve the proposed models. Extensive experiments on simulated and real-world videos demonstrate the superiority of the proposed approaches over the existing state-of-the-art approaches.
Wenfei Cao, Yao Wang 0003, Jian Sun 0009, Deyu Meng, Can Yang 0002, Andrzej Cichocki, Zongben Xu
IEEE Trans. Image Process.7
2016 Robust Low-Rank Matrix Factorization Under General Mixture Noise Distributions
abstract
Many computer vision problems can be posed as learning a low-dimensional subspace from high-dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problems using L1-norm and L2-norm losses, which mainly deal with the Laplace and Gaussian noises, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as mixture of exponential power (MoEP) distributions and then proposes a penalized MoEP (PMoEP) model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture distribution is adapted from a series of preliminary superor sub-Gaussian candidates. Moreover, by facilitating the local continuity of noise components, we embed Markov random field into the PMoEP model and then propose the PMoEP-MRF model. A generalized expectation maximization (GEM) algorithm and a variational GEM algorithm are designed to infer all parameters involved in the proposed PMoEP and the PMoEPMRF model, respectively. The superiority of our methods is demonstrated by extensive experiments on synthetic data, face modeling, hyperspectral image denoising, and background subtraction.
Xiangyong Cao, Qian Zhao 0002, Deyu Meng, Yang Chen 0057, Zongben Xu
IEEE Trans. Image Process.5
2015 Self-Paced Learning for Matrix Factorization
abstract
Matrix factorization (MF) has been attracting much attention due to its wide applications. However, since MF models are generally non-convex, most of the existing methods are easily stuck into bad local minima, especially in the presence of outliers and missing data. To alleviate this deficiency, in this study we present a new MF learning methodology by gradually including matrix elements into MF training from easy to complex. This corresponds to a recently proposed learning fashion called self-paced learning (SPL), which has been demonstrated to be beneficial in avoiding bad local minima. We also generalize the conventional binary (hard) weighting scheme for SPL to a more effective real-valued (soft) weighting manner. The effectiveness of the proposed self-paced MF method is substantiated by a series of experiments on synthetic, structure from motion and background subtraction data.
Qian Zhao 0002, Deyu Meng, Lu Jiang 0004, Qi Xie 0002, Zongben Xu, Alex Hauptmann 0001
AAAI5
2015 Learning a convolutional neural network for non-uniform motion blur removal
abstract
In this paper, we address the problem of estimating and removing non-uniform motion blur from a single blurry image. We propose a deep learning approach to predicting the probabilistic distribution of motion blur at the patch level using a convolutional neural network (CNN). We further extend the candidate set of motion kernels predicted by the CNN using carefully designed image rotations. A Markov random field model is then used to infer a dense non-uniform motion blur field enforcing motion smoothness. Finally, motion blur is removed by a non-uniform deblurring model using patch-level image prior. Experimental evaluations show that our approach can effectively estimate and remove complex non-uniform motion blur that is not handled well by previous approaches.
Jian Sun 0009, Wenfei Cao, Zongben Xu, Jean Ponce
CVPR3
2015 Low-Rank Matrix Factorization under General Mixture Noise Distributions
abstract
Many computer vision problems can be posed as learning a low-dimensional subspace from high dimensional data. The low rank matrix factorization (LRMF) represents a commonly utilized subspace learning strategy. Most of the current LRMF techniques are constructed on the optimization problem using L_1 norm and L_2 norm, which mainly deal with Laplacian and Gaussian noise, respectively. To make LRMF capable of adapting more complex noise, this paper proposes a new LRMF model by assuming noise as Mixture of Exponential Power (MoEP) distributions and proposes a penalized MoEP model by combining the penalized likelihood method with MoEP distributions. Such setting facilitates the learned LRMF model capable of automatically fitting the real noise through MoEP distributions. Each component in this mixture is adapted from a series of preliminary super-or sub-Gaussian candidates. An Expectation Maximization (EM) algorithm is also designed to infer the parameters involved in the proposed PMoEP model. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and hyperspectral image restoration.
Xiangyong Cao, Yang Chen 0057, Qian Zhao 0002, Deyu Meng, Yao Wang 0003, Zongben Xu
ICCV7
2015 A Novel Sparsity Measure for Tensor Recovery
abstract
In this paper, we propose a new sparsity regularizer for measuring the low-rank structure underneath a tensor. The proposed sparsity measure has a natural physical meaning which is intrinsically the size of the fundamental Kronecker basis to express the tensor. By embedding the sparsity measure into the tensor completion and tensor robust PCA frameworks, we formulate new models to enhance their capability in tensor recovery. Through introducing relaxation forms of the proposed sparsity measure, we also adopt the alternating direction method of multipliers (ADMM) for solving the proposed models. Experiments implemented on synthetic and multispectral image data sets substantiate the effectiveness of the proposed methods.
Qian Zhao 0002, Deyu Meng, Xu Kong, Qi Xie 0002, Wenfei Cao, Yao Wang 0003, Zongben Xu
ICCV7
2015 Robust low-rank tensor factorization by cyclic weighted median
Deyu Meng, Biao Zhang 0005, Zongben Xu, Lei Zhang 0006, Chenqiang Gao
Sci. China Inf. Sci.3
2015 Folded-concave penalization approaches to tensor completion
Wenfei Cao, Yao Wang 0003, Can Yang 0002, Xiangyu Chang, Zhi Han, Zongben Xu
Neurocomputing6
2015 A block coordinate descent approach for sparse principal component analysis
Qian Zhao 0002, Deyu Meng, Zongben Xu, Chenqiang Gao
Neurocomputing3
2015 Jackson-type inequalities for spherical neural networks with doubling weights
Shaobo Lin, Jinshan Zeng, Lin Xu 0001, Zongben Xu
Neural Networks4
2015 Error Estimate for Spherical Neural Networks Interpolation
Shaobo Lin, Jinshan Zeng, Zongben Xu
Neural Process. Lett.3
2015 The Generalization Ability of SVM Classification Based on Markov Sampling
abstract
UNLABELLED: The previously known works studying the generalization ability of support vector machine classification (SVMC) algorithm are usually based on the assumption of independent and identically distributed samples. In this paper, we go far beyond this classical framework by studying the generalization ability of SVMC based on uniformly ergodic Markov chain (u.e.M.c.) samples. We analyze the excess misclassification error of SVMC based on u.e.M.c. samples, and obtain the optimal learning rate of SVMC for u.e.M.c. SAMPLES: We also introduce a new Markov sampling algorithm for SVMC to generate u.e.M.c. samples from given dataset, and present the numerical studies on the learning performance of SVMC based on Markov sampling for benchmark datasets. The numerical studies show that the SVMC based on Markov sampling not only has better generalization ability as the number of training samples are bigger, but also the classifiers based on Markov sampling are sparsity when the size of dataset is bigger with regard to the input dimension.
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009, Baochang Zhang 0001
IEEE Trans. Cybern.4
2015 Color Image Denoising via Discriminatively Learned Iterative Shrinkage
abstract
In this paper, we propose a novel model, a discriminatively learned iterative shrinkage (DLIS) model, for color image denoising. The DLIS is a generalization of wavelet shrinkage by iteratively performing shrinkage over patch groups and whole image aggregation. We discriminatively learn the shrinkage functions and basis from the training pairs of noisy/noise-free images, which can adaptively handle different noise characteristics in luminance/chrominance channels, and the unknown structured noise in real-captured color images. Furthermore, to remove the splotchy real color noises, we design a Laplacian pyramid-based denoising framework to progressively recover the clean image from the coarsest scale to the finest scale by the DLIS model learned from the real color noises. Experiments show that our proposed approach can achieve the state-of-the-art denoising results on both synthetic denoising benchmark and real-captured color images.
Jian Sun 0009, Jian Sun 0001, Zongben Xu
IEEE Trans. Image Process.3
2015 Is Extreme Learning Machine Feasible? A Theoretical Assessment (Part II)
abstract
An extreme learning machine (ELM) can be regarded as a two-stage feed-forward neural network (FNN) learning system that randomly assigns the connections with and within hidden neurons in the first stage and tunes the connections with output neurons in the second stage. Therefore, ELM training is essentially a linear learning problem, which significantly reduces the computational burden. Numerous applications show that such a computation burden reduction does not degrade the generalization capability. It has, however, been open that whether this is true in theory. The aim of this paper is to study the theoretical feasibility of ELM by analyzing the pros and cons of ELM. In the previous part of this topic, we pointed out that via appropriately selected activation functions, ELM does not degrade the generalization capability in the sense of expectation. In this paper, we launch the study in a different direction and show that the randomness of ELM also leads to certain negative consequences. On one hand, we find that the randomness causes an additional uncertainty problem of ELM, both in approximation and learning. On the other hand, we theoretically justify that there also exist activation functions such that the corresponding ELM degrades the generalization capability. In particular, we prove that the generalization capability of ELM with Gaussian kernel is essentially worse than that of FNN with Gaussian kernel. To facilitate the use of ELM, we also provide a remedy to such a degradation. We find that the well-developed coefficient regularization technique can essentially improve the generalization capability. The obtained results reveal the essential characteristic of ELM in a certain sense and give theoretical guidance concerning how to use ELM.
Shaobo Lin, Jian Fang 0001, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2015 Is Extreme Learning Machine Feasible? A Theoretical Assessment (Part I)
abstract
An extreme learning machine (ELM) is a feedforward neural network (FNN) like learning system whose connections with output neurons are adjustable, while the connections with and within hidden neurons are randomly fixed. Numerous applications have demonstrated the feasibility and high efficiency of ELM-like systems. It has, however, been open if this is true for any general applications. In this two-part paper, we conduct a comprehensive feasibility analysis of ELM. In Part I, we provide an answer to the question by theoretically justifying the following: 1) for some suitable activation functions, such as polynomials, Nadaraya-Watson and sigmoid functions, the ELM-like systems can attain the theoretical generalization bound of the FNNs with all connections adjusted, i.e., they do not degrade the generalization capability of the FNNs even when the connections with and within hidden neurons are randomly fixed; 2) the number of hidden neurons needed for an ELM-like system to achieve the theoretical bound can be estimated; and 3) whenever the activation function is taken as polynomial, the deduced hidden layer output matrix is of full column-rank, therefore the generalized inverse technique can be efficiently applied to yield the solution of an ELM-like system, and, furthermore, for the nonpolynomial case, the Tikhonov regularization can be applied to guarantee the weak regularity while not sacrificing the generalization capability. In Part II, however, we reveal a different aspect of the feasibility of ELM: there also exists some activation functions, which makes the corresponding ELM degrade the generalization capability. The obtained results underlie the feasibility and efficiency of ELM-like systems, and yield various generalizations and improvements of the systems as well.
Shaobo Lin, Jian Fang 0001, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2015 The Generalization Ability of Online SVM Classification Based on Markov Sampling
abstract
In this paper, we consider online support vector machine (SVM) classification learning algorithms with uniformly ergodic Markov chain (u.e.M.c.) samples. We establish the bound on the misclassification error of an online SVM classification algorithm with u.e.M.c. samples based on reproducing kernel Hilbert spaces and obtain a satisfactory convergence rate. We also introduce a novel online SVM classification algorithm based on Markov sampling, and present the numerical studies on the learning ability of online SVM classification based on Markov sampling for benchmark repository. The numerical studies show that the learning performance of the online SVM classification algorithm based on Markov sampling is better than that of classical online SVM classification based on random sampling as the size of training samples is larger.
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009
IEEE Trans. Neural Networks Learn. Syst.4
2015 L1-Norm Low-Rank Matrix Factorization by Variational Bayesian Method
abstract
The L1 -norm low-rank matrix factorization (LRMF) has been attracting much attention due to its wide applications to computer vision and pattern recognition. In this paper, we construct a new hierarchical Bayesian generative model for the L1 -norm LRMF problem and design a mean-field variational method to automatically infer all the parameters involved in the model by closed-form equations. The variational Bayesian inference in the proposed method can be understood as solving a weighted LRMF problem with different weights on matrix elements based on their significance and with L2 -regularization penalties on parameters. Throughout the inference process of our method, the weights imposed on the matrix elements can be adaptively fitted so that the adverse influence of noises and outliers embedded in data can be largely suppressed, and the parameters can be appropriately regularized so that the generalization capability of the problem can be statistically guaranteed. The robustness and the efficiency of the proposed method are substantiated by a series of synthetic and real data experiments, as compared with the state-of-the-art L1 -norm LRMF methods. Especially, attributed to the intrinsic generalization capability of the Bayesian methodology, our method can always predict better on the unobserved ground truth data than existing methods.
Qian Zhao 0002, Deyu Meng, Zongben Xu, Wangmeng Zuo, Yan Yan 0002
IEEE Trans. Neural Networks Learn. Syst.3
2014 Decomposable Nonlocal Tensor Dictionary Learning for Multispectral Image Denoising
abstract
As compared to the conventional RGB or gray-scale images, multispectral images (MSI) can deliver more faithful representation for real scenes, and enhance the performance of many computer vision tasks. In practice, however, an MSI is always corrupted by various noises. In this paper we propose an effective MSI denoising approach by combinatorially considering two intrinsic characteristics underlying an MSI: the nonlocal similarity over space and the global correlation across spectrum. In specific, by explicitly considering spatial self-similarity of an MSI we construct a nonlocal tensor dictionary learning model with a group-block-sparsity constraint, which makes similar full-band patches (FBP) share the same atoms from the spatial and spectral dictionaries. Furthermore, through exploiting spectral correlation of an MSI and assuming over-redundancy of dictionaries, the constrained nonlocal MSI dictionary learning model can be decomposed into a series of unconstrained low-rank tensor approximation problems, which can be readily solved by off-the-shelf higher order statistics. Experimental results show that our method outperforms all state-of-the-art MSI denoising methods under comprehensive quantitative performance measures.
Deyu Meng, Zongben Xu, Chenqiang Gao, Yi Yang 0001, Biao Zhang 0005
CVPR3
2014 Robust Principal Component Analysis with Complex Noise
abstract
The research on robust principal component analysis (RPCA) has been attracting much attention recently. The original RPCA model assumes sparse noise, and use the L_1-norm to characterize the error term. In practice, however, the noise is much more complex and it is not appropriate to simply use a certain L_p-norm for noise modeling. We propose a generative RPCA model under the Bayesian framework by modeling data noise as a mixture of Gaussians (MoG). The MoG is a universal approximator to continuous distributions and thus our model is able to fit a wide range of noises such as Laplacian, Gaussian, sparse noises and any combinations of them. A variational Bayes algorithm is presented to infer the posterior of the proposed model. All involved parameters can be recursively updated in closed form. The advantage of our method is demonstrated by extensive experiments on synthetic data, face modeling and background subtraction.
Qian Zhao 0002, Deyu Meng, Zongben Xu, Wangmeng Zuo, Lei Zhang 0006
ICML3
2014 Hierarchical clustering driven by cognitive features
Chun-Zhong Li, Zongben Xu, Chen Qiao, Tao Luo 0006
Sci. China Inf. Sci.2
2014 The number of spanning trees in a new lexicographic product of graphs
Feng Li 0057, Zongben Xu
Sci. China Inf. Sci.3
2014 Robust sparse principal component analysis
Qian Zhao 0002, Deyu Meng, Zongben Xu
Sci. China Inf. Sci.3
2014 Two soft-thresholding based iterative algorithms for image deblurring
Jie Huang 0005, Ting-Zhu Huang, Xi-Le Zhao, Zongben Xu, Xiao-Guang Lv
Inf. Sci.4
2014 An active contour model and its algorithms with local and global Gaussian distribution fitting energies
Ting-Zhu Huang, Zongben Xu
Inf. Sci.3
2014 Almost optimal estimates for approximation and learning by radial basis function networks
Shaobo Lin, Yuanhua Rong, Zongben Xu
Mach. Learn.4
2014 Learning Rates of lq Coefficient Regularization Learning with Gaussian Kernel
abstract
Regularization is a well-recognized powerful strategy to improve the performance of a learning machine and l(q) regularization schemes with 0 < q < ∞ are central in use. It is known that different q leads to different properties of the deduced estimators, say, l(2) regularization leads to a smooth estimator, while l(1) regularization leads to a sparse estimator. Then how the generalization capability of l(q) regularization learning varies with q is worthy of investigation. In this letter, we study this problem in the framework of statistical learning theory. Our main results show that implementing l(q) coefficient regularization schemes in the sample-dependent hypothesis space associated with a gaussian kernel can attain the same almost optimal learning rates for all 0 < q < ∞. That is, the upper and lower bounds of learning rates for l(q) regularization learning are asymptotically identical for all 0 < q < ∞. Our finding tentatively reveals that in some modeling contexts, the choice of q might not have a strong impact on the generalization capability. From this perspective, q can be arbitrarily specified, or specified merely by other nongeneralization criteria like smoothness, computational complexity or sparsity.
Shaobo Lin, Jinshan Zeng, Jian Fang 0001, Zongben Xu
Neural Comput.4
2014 Generalization performance of Gaussian kernels SVMC based on Markov sampling
Jie Xu 0006, Yuan Yan Tang, Bin Zou 0002, Zongben Xu, Luoqing Li, Yang Lu 0009
Neural Networks4
2014 Restricted p-isometry properties of nonconvex block-sparse compressed sensing
Yao Wang 0003, Jianjun Wang 0003, Zongben Xu
Signal Process.3
2014 Sparse solution of underdetermined linear equations via adaptively iterative thresholding
Jinshan Zeng, Shaobo Lin, Zongben Xu
Signal Process.3
2014 A novel L1/2 regularization shooting method for Cox's proportional hazards model
Xin-Ze Luan, Yong Liang 0001, Kwong-Sak Leung, Tak-Ming Chan, Zongben Xu, Hai Zhang 0001
Soft Comput.6
2014 Sparse Bayesian Hierarchical Prior Modeling Based Cooperative Spectrum Sensing in Wideband Cognitive Radio Networks
abstract
This letter proposes a new method for cooperative spectrum sensing by exploiting sparsity. The novel scheme uses the theory of Bayesian hierarchical prior modeling in the framework of sparse Bayesian learning. This model has sparsity-inducing penalization terms leading to sparser solutions compared with typically${l_1}$norm based ones. Based on the factor graph that represents the signal model of the hierarchical prior models, the variational message passing (VMP) algorithm is implemented to estimate the power spectral density (PSD) map.
Feng Li 0057, Zongben Xu
IEEE Signal Process. Lett.2
2014 The Generalization Performance of Regularized Regression Algorithms Based on Markov Sampling
abstract
This paper considers the generalization ability of two regularized regression algorithms [least square regularized regression (LSRR) and support vector machine regression (SVMR)] based on non-independent and identically distributed (non-i.i.d.) samples. Different from the previously known works for non-i.i.d. samples, in this paper, we research the generalization bounds of two regularized regression algorithms based on uniformly ergodic Markov chain (u.e.M.c.) samples. Inspired by the idea from Markov chain Monto Carlo (MCMC) methods, we also introduce a new Markov sampling algorithm for regression to generate u.e.M.c. samples from a given dataset, and then, we present the numerical studies on the learning performance of LSRR and SVMR based on Markov sampling, respectively. The experimental results show that LSRR and SVMR based on Markov sampling can present obviously smaller mean square errors and smaller variances compared to random sampling.
Bin Zou 0002, Yuan Yan Tang, Zongben Xu, Luoqing Li, Jie Xu 0006, Yang Lu 0009
IEEE Trans. Cybern.3
2014 Spatial and Spectral Image Fusion Using Sparse Matrix Factorization
abstract
In this paper, we present a novel spatial and spectral fusion model (SASFM) that uses sparse matrix factorization to fuse remote sensing imagery with different spatial and spectral properties. By combining the spectral information from sensors with low spatial resolution (LSaR) but high spectral resolution (HSeR) (hereafter called HSeR sensors), with the spatial information from sensors with high spatial resolution (HSaR) but low spectral resolution (LSeR) (hereafter called HSaR sensors), the SASFM can generate synthetic remote sensing data with both HSaR and HSeR. Given two reasonable assumptions, the proposed model can integrate the LSaR and HSaR data via two stages. In the first stage, the model learns from the LSaR data a spectral dictionary containing pure signatures, and in the second stage, the desired HSaR and HSeR data are predicted using the learned spectral dictionary and the known HSaR data. The SASFM is tested with both simulated data and actual Landsat 7 Enhanced Thematic Mapper Plus (ETM+) and Terra Moderate Resolution Imaging Spectroradiometer (MODIS) acquisitions, and it is also compared to other representative algorithms. The experimental results demonstrate that the SASFM outperforms other algorithms in generating fused imagery with both the well-preserved spectral properties of MODIS and the spatial properties of ETM+. Generated imagery with simultaneous HSaR and HSeR opens new avenues for applications of MODIS and ETM+.
Bo Huang 0001, Hengbin Cui, Zongben Xu
IEEE Trans. Geosci. Remote. Sens.5
2014 Enhancing Low-Rank Subspace Clustering by Manifold Regularization
abstract
Recently, low-rank representation (LRR) method has achieved great success in subspace clustering (SC), which aims to cluster the data points that lie in a union of low-dimensional subspace. Given a set of data points, LRR seeks the lowest rank representation among the many possible linear combinations of the bases in a given dictionary or in terms of the data itself. However, LRR only considers the global Euclidean structure, while the local manifold structure, which is often important for many real applications, is ignored. In this paper, to exploit the local manifold structure of the data, a manifold regularization characterized by a Laplacian graph has been incorporated into LRR, leading to our proposed Laplacian regularized LRR (LapLRR). An efficient optimization procedure, which is based on alternating direction method of multipliers (ADMM), is developed for LapLRR. Experimental results on synthetic and real data sets are presented to demonstrate that the performance of LRR has been enhanced by using the manifold regularization.
Junmin Liu, Jiangshe Zhang 0001, Zongben Xu
IEEE Trans. Image Process.4
2014 Detection and Reconstruction of an Implicit Boundary Surface by Adaptively Expanding A Small Surface Patch in a 3D Image
abstract
In this paper we propose a novel and easy to use 3D reconstruction method. With the method, users only need to specify a small boundary surface patch in a 2D section image, and then an entire continuous implicit boundary surface (CIBS) can be automatically reconstructed from a 3D image. In the method, a hierarchical tracing strategy is used to grow the known boundary surface patch gradually in the 3D image. An adaptive detection technique is applied to detect boundary surface patches from different local regions. The technique is based on both context dependence and adaptive contrast detection as in the human vision system. A recognition technique is used to distinguish true boundary surface patches from the false ones in different cubes. By integrating these different approaches, a high-resolution CIBS model can be automatically reconstructed by adaptively expanding the small boundary surface patch in the 3D image. The effectiveness of our method is demonstrated by its applications to a variety of real 3D images, where the CIBS with complex shapes/branches and with varying gray values/gradient magnitudes can be well reconstructed. Our method is easy to use, which provides a valuable tool for 3D image visualization and analysis as needed in many applications.
Lisheng Wang, Pai Wang 0002, Liuhang Cheng, Shenzhi Wu, Yu-Ping Wang 0002, Zongben Xu
IEEE Trans. Vis. Comput. Graph.7
2013 A Cyclic Weighted Median Method for L1 Low-Rank Matrix Factorization with Missing Entries
abstract
A challenging problem in machine learning, information retrieval and computer vision research is how to recover a low-rank representation of the given data in the presence of outliers and missing entries. The L1-norm low-rank matrix factorization (LRMF) has been a popular approach to solving this problem. However, L1-norm LRMF is difficult to achieve due to its non-convexity and non-smoothness, and existing methods are often inefficient and fail to converge to a desired solution. In this paper we propose a novel cyclic weighted median (CWM) method, which is intrinsically a coordinate decent algorithm, for L1-norm LRMF. The CWM method minimizes the objective by solving a sequence of scalar minimization sub-problems, each of which is convex and can be easily solved by the weighted median filter. The extensive experimental results validate that the CWM method outperforms state-of-the-arts in terms of both accuracy and computational efficiency.
Deyu Meng, Zongben Xu, Lei Zhang 0006, Ji Zhao 0001
AAAI2
2013 Sparse K-Means with the l_q(0leq q< 1) Constraint for High-Dimensional Data Clustering
abstract
Sparse clustering, which aims at finding a proper partition of extremely high dimensional data set with fewest relevant features, has been attracted more and more attention. Most researches model the problem through minimizing weighted feature contributions subject to a l1constraint. However, the l0constraint is the essential constraint for sparse modeling while the l1constraint is only a convex relaxation of it. In this article, we bridge the gap between the l0constraint and the l1constraint through development of two new sparse clustering models, which are the sparse k-means with the lq(00constraint. By proving the certain forms of the optimal solution of particular lq(0 = qq(0 = q1constraint.
Yu Wang 0075, Xiangyu Chang, Rongjian Li, Zongben Xu
ICDM4
2013 Sparse logistic regression with a L1/2 penalty for gene selection in cancer classification
abstract
BACKGROUND: Microarray technology is widely used in cancer diagnosis. Successfully identifying gene biomarkers will significantly help to classify different cancer types and improve the prediction accuracy. The regularization approach is one of the effective methods for gene selection in microarray data, which generally contain a large number of genes and have a small number of samples. In recent years, various approaches have been developed for gene selection of microarray data. Generally, they are divided into three categories: filter, wrapper and embedded methods. Regularization methods are an important embedded technique and perform both continuous shrinkage and automatic gene selection simultaneously. Recently, there is growing interest in applying the regularization techniques in gene selection. The popular regularization technique is Lasso (L1), and many L1 type regularization terms have been proposed in the recent years. Theoretically, the Lq type regularization with the lower value of q would lead to better solutions with more sparsity. Moreover, the L1/2 regularization can be taken as a representative of Lq (0 <q < 1) regularizations and has been demonstrated many attractive properties. RESULTS: In this work, we investigate a sparse logistic regression with the L1/2 penalty for gene selection in cancer classification problems, and propose a coordinate descent algorithm with a new univariate half thresholding operator to solve the L1/2 penalized logistic regression. Experimental results on artificial and microarray data demonstrate the effectiveness of our proposed approach compared with other regularization methods. Especially, for 4 publicly available gene expression datasets, the L1/2 regularization method achieved its success using only about 2 to 14 predictors (genes), compared to about 6 to 38 genes for ordinary L1 and elastic net regularization approaches. CONCLUSIONS: From our evaluations, it is clear that the sparse logistic regression with the L1/2 penalty achieves higher classification accuracy than those of ordinary L1 and elastic net regularization approaches, while fewer but informative genes are selected. This is an important consideration for screening and diagnostic applications, where the goal is often to develop an accurate test using as few features as possible in order to control cost. Therefore, the sparse logistic regression with the L1/2 penalty is effective technique for gene selection in real classification problems.
Yong Liang 0001, Xin-Ze Luan, Kwong-Sak Leung, Tak-Ming Chan, Zongben Xu, Hai Zhang 0001
BMC Bioinform.6
2013 Image restoration with shifting reflective boundary conditions
Jie Huang 0005, Ting-Zhu Huang, Xi-Le Zhao, Zongben Xu
Sci. China Inf. Sci.4
2013 The learning performance of support vector machine classification based on Markov sampling
Bin Zou 0002, Zhiming Peng, Zongben Xu
Sci. China Inf. Sci.3
2013 Following the entire solution path of sparse principal component analysis by coordinate-pairwise algorithm
Deyu Meng, Hengbin Cui, Zongben Xu, Kaili Jing
Data Knowl. Eng.3
2013 Learning dictionary from signals under global sparsity constraint
Deyu Meng, Qian Zhao 0002, Yee Leung, Zongben Xu
Neurocomputing4
2013 The UPPAM continuous-time RNN model and its critical dynamics study
Chen Qiao, Wenfeng Jing, Zongben Xu
Neurocomputing3
2013 The strong convergence of visual classification method and its applications
Deyu Meng, Yee Leung, Zongben Xu
Inf. Sci.3
2013 Fast image deconvolution using closed-form thresholding formulas of regularization
Wenfei Cao, Jian Sun 0009, Zongben Xu
J. Vis. Commun. Image Represent.3
2013 Passage method for nonlinear dimensionality reduction of data on multi-cluster manifolds
Deyu Meng, Yee Leung, Zongben Xu
Pattern Recognit.3
2013 A heuristic hierarchical clustering based on multiple similarity measurements
Chun-Zhong Li, Zongben Xu, Tao Luo 0006
Pattern Recognit. Lett.2
2013 Accelerated L1/2 regularization based SAR imaging via BCR and reduced Newton skills
Jinshan Zeng, Zongben Xu, Bingchen Zhang, Wen Hong, Yirong Wu
Signal Process.2
2013 Dynamic Extreme Learning Machine and Its Approximation Capability
abstract
Extreme learning machines (ELMs) have been proposed for generalized single-hidden-layer feedforward networks which need not be neuron alike and perform well in both regression and classification applications. The problem of determining the suitable network architectures is recognized to be crucial in the successful application of ELMs. This paper first proposes a dynamic ELM (D-ELM) where the hidden nodes can be recruited or deleted dynamically according to their significance to network performance, so that not only the parameters can be adjusted but also the architecture can be self-adapted simultaneously. Then, this paper proves in theory that such D-ELM using Lebesgue p-integrable hidden activation functions can approximate any Lebesgue p-integrable function on a compact input set. Simulation results obtained over various test problems demonstrate and verify that the proposed D-ELM does a good job reducing the network size while preserving good generalization performance.
Rui Zhang 0005, Yuan Lan, Guang-Bin Huang, Zongben Xu, Yeng Chai Soh
IEEE Trans. Cybern.4
2013 Detecting Intrinsic Loops Underlying Data Manifold
abstract
Detecting intrinsic loop structures of a data manifold is the necessary prestep for the proper employment of the manifold learning techniques and of fundamental importance in the discovery of the essential representational features underlying the data lying on the loopy manifold. An effective strategy is proposed to solve this problem in this study. In line with our intuition, a formal definition of a loop residing on a manifold is first given. Based on this definition, theoretical properties of loopy manifolds are rigorously derived. In particular, a necessary and sufficient condition for detecting essential loops of a manifold is derived. An effective algorithm for loop detection is then constructed. The soundness of the proposed theory and algorithm is validated by a series of experiments performed on synthetic and real-life data sets. In each of the experiments, the essential loops underlying the data manifold can be properly detected, and the intrinsic representational features of the data manifold can be revealed along the loop structure so detected. Particularly, some of these features can hardly be discovered by the conventional manifold learning methods.
Deyu Meng, Yee Leung, Zongben Xu
IEEE Trans. Knowl. Data Eng.3
2013 Learning Capability of Relaxed Greedy Algorithms
abstract
In the practice of machine learning, one often encounters problems in which noisy data are abundant while the learning targets are imprecise and elusive. To these challenges, most of the traditional learning algorithms employ hypothesis spaces of large capacity. This has inevitably led to high computational burdens and caused considerable machine sluggishness. Utilizing greedy algorithms in this kind of learning environment has greatly improved machine performance. The best existing learning rate of various greedy algorithms is proved to achieve the order of (m/log m)(-1/2), where m is the sample size. In this paper, we provide a relaxed greedy algorithm and study its learning capability. We prove that the learning rate of the new relaxed greedy algorithm is faster than the order m(-1/2). Unlike many other greedy algorithms, which are often indecisive issuing a stopping order to the iteration process, our algorithm has a clearly established stopping criteria.
Shaobo Lin, Yuanhua Rong, Xingping Sun, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2013 Generalization Performance of Fisher Linear Discriminant Based on Markov Sampling
abstract
Fisher linear discriminant (FLD) is a well-known method for dimensionality reduction and classification that projects high-dimensional data onto a low-dimensional space where the data achieves maximum class separability. The previous works describing the generalization ability of FLD have usually been based on the assumption of independent and identically distributed (i.i.d.) samples. In this paper, we go far beyond this classical framework by studying the generalization ability of FLD based on Markov sampling. We first establish the bounds on the generalization performance of FLD based on uniformly ergodic Markov chain (u.e.M.c.) samples, and prove that FLD based on u.e.M.c. samples is consistent. By following the enlightening idea from Markov chain Monto Carlo methods, we also introduce a Markov sampling algorithm for FLD to generate u.e.M.c. samples from a given data of finite size. Through simulation studies and numerical studies on benchmark repository using FLD, we find that FLD based on u.e.M.c. samples generated by Markov sampling can provide smaller misclassification rates compared to i.i.d. samples.
Bin Zou 0002, Luoqing Li, Zongben Xu, Tao Luo 0006, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.3
2012 SAR range ambiguity suppression via sparse regularization
abstract
Range ambiguity in synthetic aperture radar (SAR) imaging primarily arises from scattered energy of bright targets outside the interested region. So to reduce the ambiguity, we need to identify these targets additionally, which yields an ill-posed problem. To find a feasible solution where the range ambiguity can be sufficiently reduced, we propose in this paper a new method using compressed sensing, a theory which tells when sparse signal can be reconstruct from undetermined linear system, by observing that the recognizable targets are approximately sparse in the ambiguous range zones. Therefore, it is possible to reconstruct the main region and identify the ambiguous targets simultaneously. The simulation results demonstrate the validation of the proposed method.
Jian Fang 0001, Zongben Xu, Chenglong Jiang, Bingchen Zhang, Wen Hong
IGARSS2
2012 MOEA/D with Iterative Thresholding Algorithm for Sparse Optimization Problems
Hui Li 0020, Xiaolei Su, Zongben Xu, Qingfu Zhang 0001
PPSN (2)3
2012 Efficient DPCA SAR imaging with fast iterative spectrum reconstruction method
Jian Fang 0001, Jinshan Zeng, Zongben Xu
Sci. China Inf. Sci.3
2012 Estimation of convergence rate for multi-regression learning algorithm
Zongben Xu, Feilong Cao
Sci. China Inf. Sci.1
2012 Sparse SAR imaging based on L 1/2 regularization
Jinshan Zeng, Jian Fang 0001, Zongben Xu
Sci. China Inf. Sci.3
2012 The essential ability of sparse reconstruction of different compressive sensing strategies
Hai Zhang 0001, Yong Liang 0001, HaiLiang Gou, Zongben Xu
Sci. China Inf. Sci.4
2012 Critical dynamics study on recurrent neural networks: Globally exponential stability
Chen Qiao, Zongben Xu
Neurocomputing2
2012 Kronecker product approximations for image restoration with whole-sample symmetric boundary conditions
Xiao-Guang Lv, Ting-Zhu Huang, Zongben Xu, Xi-Le Zhao
Inf. Sci.3
2012 Improve robustness of sparse PCA by L1-norm maximization
Deyu Meng, Qian Zhao 0002, Zongben Xu
Pattern Recognit.3
2012 L1/2 Regularization: A Thresholding Representation Theory and a Fast Solver
abstract
The special importance of L1/2 regularization has been recognized in recent studies on sparse modeling (particularly on compressed sensing). The L1/2 regularization, however, leads to a nonconvex, nonsmooth, and non-Lipschitz optimization problem that is difficult to solve fast and efficiently. In this paper, through developing a threshoding representation theory for L1/2 regularization, we propose an iterative half thresholding algorithm for fast solution of L1/2 regularization, corresponding to the well-known iterative soft thresholding algorithm for L1 regularization, and the iterative hard thresholding algorithm for L0 regularization. We prove the existence of the resolvent of gradient of ||x||1/2(1/2), calculate its analytic expression, and establish an alternative feature theorem on solutions of L1/2 regularization, based on which a thresholding representation of solutions of L1/2 regularization is derived and an optimal regularization parameter setting rule is formulated. The developed theory provides a successful practice of extension of the well- known Moreau's proximity forward-backward splitting theory to the L1/2 regularization case. We verify the convergence of the iterative half thresholding algorithm and provide a series of experiments to assess performance of the algorithm. The experiments show that the half algorithm is effective, efficient, and can be accepted as a fast solver for L1/2 regularization. With the new algorithm, we conduct a phase diagram study to further demonstrate the superiority of L1/2 regularization over L1 regularization.
Zongben Xu, Xiangyu Chang, Fengmin Xu, Hai Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2012 Universal Approximation of Extreme Learning Machine With Adaptive Growth of Hidden Nodes
abstract
Extreme learning machines (ELMs) have been proposed for generalized single-hidden-layer feedforward networks which need not be neuron-like and perform well in both regression and classification applications. In this brief, we propose an ELM with adaptive growth of hidden nodes (AG-ELM), which provides a new approach for the automated design of networks. Different from other incremental ELMs (I-ELMs) whose existing hidden nodes are frozen when the new hidden nodes are added one by one, in AG-ELM the number of hidden nodes is determined in an adaptive way in the sense that the existing networks may be replaced by newly generated networks which have fewer hidden nodes and better generalization performance. We then prove that such an AG-ELM using Lebesgue p-integrable hidden activation functions can approximate any Lebesgue p-integrable function on a compact input set. Simulation results demonstrate and verify that this new approach can achieve a more compact network architecture than the I-ELM.
Rui Zhang 0005, Yuan Lan, Guang-Bin Huang, Zongben Xu
IEEE Trans. Neural Networks Learn. Syst.4
2012 Global Convergence of Online BP Training With Dynamic Learning Rate
abstract
The online backpropagation (BP) training procedure has been extensively explored in scientific research and engineering applications. One of the main factors affecting the performance of the online BP training is the learning rate. This paper proposes a new dynamic learning rate which is based on the estimate of the minimum error. The global convergence theory of the online BP training procedure with the proposed learning rate is further studied. It is proved that: 1) the error sequence converges to the global minimum error; and 2) the weight sequence converges to a fixed point at which the error function attains its global minimum. The obtained global convergence theory underlies the successful applications of the online BP training procedure. Illustrative examples are provided to support the theoretical analysis.
Rui Zhang 0005, Zongben Xu, Guang-Bin Huang, Dianhui Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2011 Video Primal Sketch: A generic middle-level representation of video
abstract
This paper presents a middle-level video representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME/MRF model with spatio-temporal filters to implicitly represent textured motion, such as water and fire, by matching feature statistics, i.e. histograms. This paper makes three contributions: i) learning a dictionary of video primitives as parametric generative model; ii) studying the Spatio-Temporal FRAME (ST-FRAME) model for modeling and synthesizing textured motion; and iii) developing a parsimonious hybrid model for generic video representation. VPS selects the proper representation automatically and is compatible with high-level action representations. In the experiments, we synthesize a series of dynamic textures, reconstruct real videos and show varying VPS over the change of densities causing by the scale transition in videos.
Zhi Han, Zongben Xu, Song-Chun Zhu
ICCV2
2011 SAR imaging from compressed measurements based on L1/2 regularization
abstract
In this paper, a novel synthetic aperture radar (SAR) imaging method based on L1/2regularization is proposed. Our method implements SAR imaging from compressed measurements with high resolution, enhanced features, reduced sidelobes and suppressed artifacts. Real SAR data experiments are implemented to demonstrate the outperformance of our method. The experiment results demonstrate that our method needs far below the traditional Nyquist rate to guarantee successful imaging. Compared to the prevalent L1regularization-based methods, there is a significant reduction of the sampling rate for SAR imaging. The sampling rate used by our method is about half of the L1regularization-based methods in the real SAR data experiments.
Jinshan Zeng, Zongben Xu, Bingchen Zhang, Wen Hong, Yirong Wu
IGARSS2
2011 A new quality assessment criterion for nonlinear dimensionality reduction
Deyu Meng, Yee Leung, Zongben Xu
Neurocomputing3
2011 Estimation of learning rate of least square algorithm via Jackson operator
Feilong Cao, Zongben Xu
Neurocomputing3
2011 Incremental Alignment Manifold Learning
Zhi Han, Deyu Meng, Zongben Xu, Nannan Gu
J. Comput. Sci. Technol.3
2011 Essential rate for approximation by spherical neural networks
Shaobo Lin, Feilong Cao, Zongben Xu
Neural Networks3
2011 Gradient Profile Prior and Its Applications in Image Super-Resolution and Enhancement
abstract
In this paper, we propose a novel generic image prior-gradient profile prior, which implies the prior knowledge of natural image gradients. In this prior, the image gradients are represented by gradient profiles, which are 1-D profiles of gradient magnitudes perpendicular to image structures. We model the gradient profiles by a parametric gradient profile model. Using this model, the prior knowledge of the gradient profiles are learned from a large collection of natural images, which are called gradient profile prior. Based on this prior, we propose a gradient field transformation to constrain the gradient fields of the high resolution image and the enhanced image when performing single image super-resolution and sharpness enhancement. With this simple but very effective approach, we are able to produce state-of-the-art results. The reconstructed high resolution images or the enhanced images are sharp while have rare ringing or jaggy artifacts.
Jian Sun 0009, Jian Sun 0001, Zongben Xu, Harry Shum
IEEE Trans. Image Process.3
2010 L1/2 regularization
Zongben Xu, Hai Zhang 0001, Yao Wang 0003, Xiangyu Chang, Yong Liang 0001
Sci. China Inf. Sci.1
2010 Approximation capability of interpolation neural networks
Feilong Cao, Shaobo Lin, Zongben Xu
Neurocomputing3
2010 On the P-critical dynamics analysis of projection recurrent neural networks
Chen Qiao, Zongben Xu
Neurocomputing2
2010 New study on neural networks: the essential order of approximation
Jianjun Wang 0003, Zongben Xu
Neural Networks2
2010 Scale selection for anisotropic diffusion filter by Markov random field model
Jian Sun 0009, Zongben Xu
Pattern Recognit.2
2010 Image Inpainting by Patch Propagation Using Patch Sparsity
abstract
This paper introduces a novel examplar-based inpainting algorithm through investigating the sparsity of natural image patches. Two novel concepts of sparsity at the patch level are proposed for modeling the patch priority and patch representation, which are two crucial steps for patch propagation in the examplar-based inpainting approach. First, patch structure sparsity is designed to measure the confidence of a patch located at the image structure (e.g., the edge or corner) by the sparseness of its nonzero similarities to the neighboring patches. The patch with larger structure sparsity will be assigned higher priority for further inpainting. Second, it is assumed that the patch to be filled can be represented by the sparse linear combination of candidate patches under the local patch consistency constraint in a framework of sparse representation. Compared with the traditional examplar-based inpainting approach, structure sparsity enables better discrimination of structure and texture, and the patch sparse representation forces the newly inpainted regions to be sharp and consistent with the surrounding textures. Experiments on synthetic and natural images show the advantages of the proposed approach.
Zongben Xu, Jian Sun 0009
IEEE Trans. Image Process.1
2009 Lower estimation of approximation rate for neural networks
Feilong Cao, Zongben Xu
Sci. China Ser. F Inf. Sci.3
2009 A critical global convergence analysis of recurrent neural networks with general projection mappings
Chen Qiao, Zongben Xu
Neurocomputing2
2009 Learning from uniformly ergodic Markov chains
Bin Zou 0002, Hai Zhang 0001, Zongben Xu
J. Complex.3
2009 The generalization performance of ERM algorithm with strongly mixing observations
Bin Zou 0002, Luoqing Li, Zongben Xu
Mach. Learn.3
2009 When Does Online BP Training Converge?
abstract
The backpropogation (BP) neural networks have been widely applied in scientific research and engineering. The success of the application, however, relies upon the convergence of the training procedure involved in the neural network learning. We settle down the convergence analysis issue through proving two fundamental theorems on the convergence of the online BP training procedure. One theorem claims that under mild conditions, the gradient sequence of the error function will converge to zero (the weak convergence), and another theorem concludes the convergence of the weight sequence defined by the procedure to a fixed value at which the error function attains its minimum (the strong convergence). The weak convergence theorem sharpens and generalizes the existing convergence analysis conducted before, while the strong convergence theorem provides new analysis results on convergence of the online BP training procedure. The results obtained reveal that with any analytic sigmoid activation function, the online BP training procedure is always convergent, which then underlies successful application of the BP neural networks.
Zongben Xu, Rui Zhang 0005, Wenfeng Jing
IEEE Trans. Neural Networks1
2009 Some Characterizations of Global Exponential Stability of a Generic Class of Continuous-Time Recurrent Neural Networks
abstract
This paper reveals two important characterizations of global exponential stability (GES) of a generic class of continuous-time recurrent neural networks. First, we show that GES of the neural networks can be fully characterized by global asymptotic stability (GAS) of the networks plus the condition that the maximum abscissa of spectral set of Jacobian matrix of the neural networks at the unique equilibrium point is less than zero. This result provides a very useful and direct way to distinguish GES from GAS for the neural networks. Second, we show that when the neural networks have small state feedback coefficients, the supremum of exponential convergence rates (ECRs) of trajectories of the neural networks is exactly equal to the absolute value of the maximum abscissa of spectral set of Jacobian matrix of the neural networks at the unique equilibrium point. Here, the supremum of ECRs indicates the potentially fastest speed of trajectory convergence. The obtained results are helpful in understanding the essence of GES and clarifying the difference between GES and GAS of the continuous-time recurrent neural networks.
Lisheng Wang, Rui Zhang 0005, Zongben Xu
IEEE Trans. Syst. Man Cybern. Part B3
2009 Fast and Efficient Strategies for Model Selection of Gaussian Support Vector Machine
abstract
Two strategies for selecting the kernel parameter (sigma) and the penalty coefficient (C) of Gaussian support vector machines (SVMs) are suggested in this paper. Based on viewing the model parameter selection problem as a recognition problem in visual systems, a direct parameter setting formula for the kernel parameter is derived through finding a visual scale at which the global and local structures of the given data set can be preserved in the feature space, and the difference between the two structures can be maximized. In addition, we propose a heuristic algorithm for the selection of the penalty coefficient through identifying the classification extent of a training datum in the implementation process of the sequential minimal optimization (SMO) procedure, which is a well-developed and commonly used algorithm in SVM training. We then evaluate the suggested strategies with a series of experiments on 13 benchmark problems and three real-world data sets, as compared with the traditional 5-cross validation (5-CV) method and the recently developed radius-margin bound (RM) method. The evaluation shows that in terms of efficiency and generalization capabilities, the new strategies outperform the current methods, and the performance is uniform and stable.
Zongben Xu, Ming-Wei Dai, Deyu Meng
IEEE Trans. Syst. Man Cybern. Part B1
2008 Image super-resolution using gradient profile prior
abstract
In this paper, we propose an image super-resolution approach using a novel generic image prior - gradient profile prior, which is a parametric prior describing the shape and the sharpness of the image gradients. Using the gradient profile prior learned from a large number of natural images, we can provide a constraint on image gradients when we estimate a hi-resolution image from a low-resolution image. With this simple but very effective prior, we are able to produce state-of-the-art results. The reconstructed hi-resolution image is sharp while has rare ringing or jaggy artifacts.
Jian Sun 0009, Zongben Xu, Harry Shum
CVPR2
2008 An Improvement to Ant Colony Optimization Heuristic
Youmei Li, Zongben Xu, Feilong Cao
ISNN (1)2
2008 The estimate for approximation error of neural networks: A constructive approach
Feilong Cao, Tingfan Xie, Zongben Xu
Neurocomputing3
2008 The essential order of approximation for Suzuki's neural networks
Feng-Jun Li, Zongben Xu, Yue-Ting Zhou
Neurocomputing2
2008 A comparative study for content-based dynamic spam classification using four machine learning algorithms
Zongben Xu
Knowl. Based Syst.2
2008 Latent semantic analysis for text categorization using neural network
Zongben Xu, Cheng-hua Li
Knowl. Based Syst.2
2008 Improving geodesic distance estimation based on locally linear assumption
Deyu Meng, Yee Leung, Zongben Xu, Tung Fung, Qingfu Zhang 0001
Pattern Recognit. Lett.3
2008 Standard Forms of Stabilizer and Normalizer Matrices for Additive Quantum Codes
abstract
In this correspondence, we use a symplectic geometry over the binary field to discuss the equivalence of additive codes over the quaternary field and the equivalence of additive quantum codes. We establish the existence of a standard form of the stabilizer and the normalizer matrices for additive quantum codes. Thus, we present the quantum analogue of standard forms of generator and parity check matrices for systematic linear codes in classical coding theory for additive quantum codes.
Ruihu Li, Zongben Xu, Xueliang Li 0001
IEEE Trans. Inf. Theory2
2008 On The Classification of Binary Optimal Self-Orthogonal Codes
abstract
The classification of binary [n,k,d] codes withdgess2k-1 and without zero coordinates is reduced to the classification of binary [(2k-1)c(k,s,t)+t,k,d] code forn=(2k-1)s+t,sges 1 and 1 lestles 2k-2, wherec(k,s,t) les min{s,t} is a function ofk,s, andt. Binary [15s+t, 4] optimal self-orthogonal codes are characterized by systems of linear equations. Based on these two results, the complete classification of [15s+t,4] optimal self-orthogonal codes fortisin {1,2,6,7,8,9,13,14} andsges 1 is obtained, and the generator matrices and weight polynomials of these 4-dimensional optimal self-orthogonal codes are also given.
Ruihu Li, Zongben Xu, Xuejun Zhao
IEEE Trans. Inf. Theory2
2008 Nonlinear Dimensionality Reduction of Data Lying on the Multicluster Manifold
abstract
A new method, which is called decomposition-composition (D-C) method, is proposed for the nonlinear dimensionality reduction (NLDR) of data lying on the multicluster manifold. The main idea is first to decompose a given data set into clusters and independently calculate the low-dimensional embeddings of each cluster by the decomposition procedure. Based on the intercluster connections, the embeddings of all clusters are then composed into their proper positions and orientations by the composition procedure. Different from other NLDR methods for multicluster data, which consider associatively the intracluster and intercluster information, the D-C method capitalizes on the separate employment of the intracluster neighborhood structures and the intercluster topologies for effective dimensionality reduction. This, on one hand, isometrically preserves the rigid-body shapes of the clusters in the embedding process and, on the other hand, guarantees the proper locations and orientations of all clusters. The theoretical arguments are supported by a series of experiments performed on the synthetic and real-life data sets. In addition, the computational complexity of the proposed method is analyzed, and its efficiency is theoretically analyzed and experimentally demonstrated. Related strategies for automatic parameter selection are also examined.
Deyu Meng, Yee Leung, Tung Fung, Zongben Xu
IEEE Trans. Syst. Man Cybern. Part B4
2007 Immunity diversity based multi-agent intrusion detection
abstract
In this paper, we propose a new method combining artificial immune with support vector machine for intrusion detection, where SVM is used as a core classification algorithm for detector. We introduce immunity diversity concept and we utilize immunity approach to create diversity detectors. We embed detector in Agents in use of the communication mechanism between the Agents, integrate each detection Agent’s result to get the judgment of intrusion detection. This distributing character makes a more robust system. Experiments show that this approach has higher detection accuracy than single SVM and Bagging.
Yu Gu 0007, Jiashu Zhao, Zongben Xu
IEEE Congress on Evolutionary Computation4
2007 Flash Cut: Foreground Extraction with Flash and No-flash Image Pairs
abstract
In this paper, we propose a novel approach for foreground layer extraction using flash/no-flash image pairs, which we call flash cut. Flash cut is based on the simple observation that only the foreground is significantly brightened by the flash and the background appearance change is very small, if the background is distant. Changes due to flash, motion, and color information are fused in an MRF framework to produce high quality segmentation results. Flash cut handles some amount of camera shake, and foreground motion, which makes it practical for anyone with a flash-equipped camera to use. We validate our approach on a variety of indoor and outdoor examples.
Jian Sun 0009, Jian Sun 0001, Sing Bing Kang, Zongben Xu, Xiaoou Tang, Harry Shum
CVPR4
2007 Neural Network Training Using Genetic Algorithm with a Novel Binary Encoding
Yong Liang 0001, Kwong-Sak Leung, Zongben Xu
ISNN (2)3
2007 New Critical Analysis on Global Convergence of Recurrent Neural Networks with Projection Mappings
Chen Qiao, Zongben Xu
ISNN (3)2
2007 Transformation of rough set models
Daowu Pei, Zongben Xu
Knowl. Based Syst.2
2006 The Essential Approximation Order for Neural Networks with Trigonometric Hidden Layer Units
Chunmei Ding, Feilong Cao, Zongben Xu
ISNN (1)3
2006 A New Pre-processing Method for Regression
Wenfeng Jing, Deyu Meng, Ming-Wei Dai, Zongben Xu
ISNN (2)4
2006 Integral Transform and Its Application to Neural Network Approximation
Feng-Jun Li, Zongben Xu
ISNN (1)2
2006 An Edge Preserving Regularization Model for Image Restoration Based on Hopfield Neural Network
Jian Sun 0009, Zongben Xu
ISNN (2)2
2006 Approximation Bound of Mixture Networks in Lomegap Spaces
Zongben Xu, Jianjun Wang 0003, Deyu Meng
ISNN (1)1
2006 The essential order of approximation for nearly exponential type neural networks
Zongben Xu, Jianjun Wang 0003
Sci. China Ser. F Inf. Sci.1
2005 Pointwise Approximation for Neural Networks
Feilong Cao, Zongben Xu, Youmei Li
ISNN (1)2
2005 Generalization and Property Analysis of GENET
Youmei Li, Zongben Xu, Feilong Cao
ISNN (1)2
2005 A New Approach for Classification: Visual Simulation Point of View
Zongben Xu, Deyu Meng, Wenfeng Jing
ISNN (2)1
2005 Simultaneous Lp-approximation order for neural networks
Zongben Xu, Fei-Long Cao
Neural Networks1
2005 Scalable model-based cluster analysis using clustering features
Huidong Jin 0001, Kwong-Sak Leung, Man Leung Wong, Zongben Xu
Pattern Recognit.4
2004 Approximation Bounds by Neural Networks in Lpomega
Jianjun Wang 0003, Zongben Xu, Weijun Xu
ISNN (1)2
2004 The essential order of approximation for neural networks
Zongben Xu, Feilong Cao
Sci. China Ser. F Inf. Sci.1
2004 Strong fuzzy compact sets and ultra-fuzzy compact sets in L-topological spaces
Sheng-Gang Li, Zongben Xu
Fuzzy Sets Syst.3
2004 An expanding self-organizing neural network for the traveling salesman problem
Kwong-Sak Leung, Huidong Jin 0001, Zongben Xu
Neurocomputing3
2004 A heuristic training for support vector regression
Zongben Xu
Neurocomputing2
2004 A comparative study of two modeling approaches in neural networks
Zongben Xu, Hong Qiao, Bo Zhang 0006
Neural Networks1
2003 Determination of the spread parameter in the Gaussian kernel for classification and regression
Zongben Xu, Weizhen Lu
Neurocomputing2
2003 An efficient self-organizing map designed by genetic algorithms for the traveling salesman problem
abstract
As a typical combinatorial optimization problem, the traveling salesman problem (TSP) has attracted extensive research interest. In this paper, we develop a self-organizing map (SOM) with a novel learning rule. It is called the integrated SOM (ISOM) since its learning rule integrates the three learning mechanisms in the SOM literature. Within a single learning step, the excited neuron is first dragged toward the input city, then pushed to the convex hull of the TSP, and finally drawn toward the middle point of its two neighboring neurons. A genetic algorithm is successfully specified to determine the elaborate coordination among the three learning mechanisms as well as the suitable parameter setting. The evolved ISOM (eISOM) is examined on three sets of TSP to demonstrate its power and efficiency. The computation complexity of the eISOM is quadratic, which is comparable to other SOM-like neural networks. Moreover, the eISOM can generate more accurate solutions than several typical approaches for TSP including the SOM developed by Budinich, the expanding SOM, the convex elastic net, and the FLEXMAP algorithm. Though its solution accuracy is not yet comparable to some sophisticated heuristics, the eISOM is one of the most accurate neural networks for the TSP.
Huidong Jin 0001, Kwong-Sak Leung, Man Leung Wong, Zongben Xu
IEEE Trans. Syst. Man Cybern. Part B4
2003 A reference model approach to stability analysis of neural networks
abstract
In this paper, a novel methodology called a reference model approach to stability analysis of neural networks is proposed. The core of the new approach is to study a neural network model with reference to other related models, so that different modeling approaches can be combinatively used and powerfully cross-fertilized. Focused on two representative neural network modeling approaches (the neuron state modeling approach and the local field modeling approach), we establish a rigorous theoretical basis on the feasibility and efficiency of the reference model approach. The new approach has been used to develop a series of new, generic stability theories for various neural network models. These results have been applied to several typical neural network systems including the Hopfield-type neural networks, the recurrent back-propagation neural networks, the BSB-type neural networks, the bound-constraints optimization neural networks, and the cellular neural networks. The results obtained unify, sharpen or generalize most of the existing stability assertions, and illustrate the feasibility and power of the new method.
Hong Qiao, Zongben Xu, Bo Zhang 0006
IEEE Trans. Syst. Man Cybern. Part B3
2002 An automata network for performing combinatorial optimization
Zongben Xu, Huidong Jin 0001, Kwong-Sak Leung, Yee Leung, Chak-Kuen Wong
Neurocomputing1
2002 The Algorithm on Knowledge Reduction in Incomplete Information Systems
Jiye Liang, Zongben Xu
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2002 Inclusion degree: a perspective on measures for rough set data analysis
Zongben Xu, Jiye Liang, Chuangyin Dang, Kwai-Sang Chin
Inf. Sci.1
2002 A new approach to stability of neural networks with time-varying delays
Hong Qiao, Zongben Xu
Neural Networks3
2001 A new model of simulated evolutionary computation-convergence analysis and specifications
abstract
There have been various algorithms designed for simulating natural evolution. This paper proposes a new simulated evolutionary computation model called the abstract evolutionary algorithm (AEA), which unifies most of the currently known evolutionary algorithms and describes the evolution as an abstract stochastic process composed of two fundamental operators: selection and evolution operators. By axiomatically characterizing the properties of the fundamental selection and evolution operators, several general convergence theorems and convergence rate estimations for the AEA are established. The established theorems are applied to a series of known evolutionary algorithms, directly fielding new convergence conditions and convergence rate estimations of various specific genetic algorithms and evolutionary strategies. The present work provides a significant step toward the establishment of a unified theory of simulated evolutionary computation.
Kwong-Sak Leung, Qihong Duan, Zongben Xu, Chak-Kuen Wong
IEEE Trans. Evol. Comput.3
2001 Nonlinear measures: a new approach to exponential stability analysis for Hopfield-type neural networks
abstract
In this paper, a new concept called nonlinear measure is introduced to quantify stability of nonlinear systems in the way similar to the matrix measure for stability of linear systems. Based on the new concept, a novel approach for stability analysis of neural networks is developed. With this approach, a series of new sufficient conditions for global and local exponential stability of Hopfield type neural networks is presented, which generalizes those existing results. By means of the introduced nonlinear measure, the exponential convergence rate of the neural networks to stable equilibrium point is estimated, and, for local stability, the attraction region of the stable equilibrium point is characterized. The developed approach can be generalized to stability analysis of other general nonlinear systems.
Hong Qiao, Zongben Xu
IEEE Trans. Neural Networks3
2001 A new data processing method based on a biological model of the compound eye: direction quantization representation
abstract
This paper presents a new data representation method called direction quantization representation (DQR) which is motivated by a simplified geometric model of biological compound eye and used in describing the shape of convex hulls of objects. Advantages of DQR include high efficiency and stability in numerical computation, convenience for semidynamic maintenance, suitability for parallel implementation, and applicability to various convex set related problems. Several practical applications are presented which show the feasibility and powerfulness of DQR.
Hong Qiao, Jiangshe Zhang 0001, Zongben Xu
IEEE Trans. Syst. Man Cybern. Part A3
2000 Characterizations of minimal T3 L-fuzzy topological spaces
Sheng-Gang Li, Zongben Xu
Inf. Sci.2
2000 Clustering by Scale-Space Filtering
abstract
In pattern recognition and image processing, the major application areas of cluster analysis, human eyes seem to possess a singular aptitude to group objects and find important structures in an efficient and effective way. Thus, a clustering algorithm simulating a visual system may solve some basic problems in these areas of research. From this point of view, we propose a new approach to data clustering by modeling the blurring effect of lateral retinal interconnections based on scale space theory. In this approach, a data set is considered as an image with each light point located at a datum position. As we blur this image, smaller light blobs merge into larger ones until the whole image becomes one light blob at a low enough level of resolution. By identifying each blob with a cluster, the blurring process generates a family of clustering along the hierarchy. The advantages of the proposed approach are: 1) The derived algorithms are computationally stable and insensitive to initialization and they are totally free from solving difficult global optimization problems. 2) It facilitates the construction of new checks on cluster validity and provides the final clustering a significant degree of robustness to noise in data and change in scale. 3) It is more robust in cases where hyperellipsoidal partitions may not be assumed. 4) it is suitable for the task of preserving the structure and integrity of the outliers in the clustering process. 5) The clustering is highly consistent with that perceived by human eyes. 6) The new approach provides a unified framework for scale-related clustering algorithms derived from many different fields such as estimation theory, recurrent signal processing on self-organization feature maps, information theory and statistical mechanics, and radial basis function neural networks.
Yee Leung, Jiangshe Zhang 0001, Zongben Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 The optimal encodings for biased association in linear associative memories
Yee Leung, Tian-Xin Dong, Zongben Xu
Neural Networks3
1998 A genetic algorithm for the multiple destination routing problems
abstract
The multiple destination routing (MDR) problem can be formulated as finding a minimal cost tree which contains designated source and multiple destination nodes so that certain constraints in a given communication network are satisfied. This is a typical NP-hard problem, and therefore only heuristic algorithms are of practical value. As a first step, a new genetic algorithm is developed to solve the MDR problems without constraints. It is based on the transformation of the underlying network of an MDR problem into its distance complete form, a natural chromosome representation of a minimal spanning tree (an individual), and a completely new computation of the fitness of individual. Compared with the known genetic algorithms and heuristic algorithms for the same problem, the proposed algorithm has several advantages. First, it guarantees convergence to an optimal solution with probability one. Second, not only are the resultant solutions all feasible, the solution quality is also much higher than that obtained by the other methods (indeed, in almost every case in our simulations, the algorithm can find the optimal solution of the problem). Third, the algorithm is of low computational complexity, and this can be decreased dramatically as the number of destination nodes in the problem increases. The simulation studies for the sparse and dense networks all demonstrate that the proposed algorithm is highly robust and very efficient in the sense of yielding high-quality solutions.
Yee Leung, Zongben Xu
IEEE Trans. Evol. Comput.3
1998 Optimal neural network algorithm for on-line string matching
abstract
We consider an online string matching problem in which we find all the occurrences of a pattern of m characters in a text of n characters, where all the characters of the pattern are available before processing, while the characters of the text are input one after the other. We propose a space-time optimal parallel algorithm for this problem using a neural network approach, This algorithm uses m McCulloch-Pitts neurons connected as a linear array. It processes every input character of the text in one step and hence it requires at most n iteration steps.
Yiu-Wing Leung, Jiangshe Zhang 0001, Zongben Xu
IEEE Trans. Syst. Man Cybern. Part B3
1997 Degree of population diversity - a perspective on premature convergence in genetic algorithms and its Markov chain analysis
abstract
In this paper, a concept of degree of population diversity is introduced to quantitatively characterize and theoretically analyze the problem of premature convergence in genetic algorithms (GAs) within the framework of Markov chain. Under the assumption that the mutation probability is zero, the search ability of GA is discussed. It is proved that the degree of population diversity converges to zero with probability one so that the search ability of a GA decreases and premature convergence occurs. Moreover, an explicit formula for the conditional probability of allele loss at a certain bit position is established to show the relationships between premature convergence and the GA parameters, such as population size, mutation probability, and some population statistics. The formula also partly answers the questions of to where a GA most likely converges. The theoretical results are all supported by the simulation experiments.
Yee Leung, Zongben Xu
IEEE Trans. Neural Networks3
1997 Neural networks for convex hull computation
abstract
Computing convex hull is one of the central problems in various applications of computational geometry. In this paper, a convex hull computing neural network (CHCNN) is developed to solve the related problems in the N-dimensional spaces. The algorithm is based on a two-layered neural network, topologically similar to ART, with a newly developed adaptive training strategy called excited learning. The CHCNN provides a parallel online and real-time processing of data which, after training, yields two closely related approximations, one from within and one from outside, of the desired convex hull. It is shown that accuracy of the approximate convex hulls obtained is around O[K(-1)(N-1/)], where K is the number of neurons in the output layer of the CHCNN. When K is taken to be sufficiently large, the CHCNN can generate any accurate approximate convex hull. We also show that an upper bound exists such that the CHCNN will yield the precise convex hull when K is larger than or equal to this bound. A series of simulations and applications is provided to demonstrate the feasibility, effectiveness, and high efficiency of the proposed algorithm.
Yee Leung, Jiangshe Zhang 0001, Zongben Xu
IEEE Trans. Neural Networks3
1996 Asymmetric Hopfield-type networks: Theory and applications
Zongben Xu, Guo-Qing Hu, Chung-Ping Kwong
Neural Networks1
1996 A Decomposition Principle for Complexity Reduction of Artificial Neural Networks
Zongben Xu, Chung-Ping Kwong
Neural Networks1
1996 Some efficient strategies for improving the eigenstructure method in synthesis of feedback neural networks
abstract
Two efficient strategies are proposed for improving the eigenstructure method from the best approximation projector point of view. Interpreted as two complementary best approximation projectors, the method is reformulated in a much more simplified form. We develop a new synthesis procedure through constructing the related best approximation projectors by using a simple recursive formula, which improves on the existing eigenstructure method not only in the significant reduction of the computational complexity but also in the incorporation of the learning capability comparable to the outer product method. The networks designed by the present procedure outperform those designed by some other known methods. We also propose a new forgetting algorithm for deleting any specific existing memories in a synthesized network. The algorithm performs efficiently and reliably, which particularly eliminates the overforgetting drawback of the Yen-Michel algorithm (1991, 1992). The feasibility and effectiveness of the algorithm are supported by theoretical analysis and computer simulations.
Zongben Xu, Guo-Qing Hu, Chung-Ping Kwong
IEEE Trans. Neural Networks1
1995 Multiple-Valued Feedback and Recurrent Correlation Neural Networks
Z.-Y. Chen, Chung-Ping Kwong, Zongben Xu
Neural Comput. Appl.3
1995 A competitive associative memory model and its dynamics
abstract
Conventional associative memory networks perform "noncompetitive recognition" or "competitive recognition in distance". In this paper a "competitive recognition" associative memory model is introduced which simulates the competitive persistence of biological species. Unlike most of the conventional networks, the proposed model takes only the prototype patterns as its equilibrium points, so that the spurious points are effectively excluded. Furthermore, it is shown that, as the competitive parameters vary, the network has a unique stable equilibrium point corresponding to the winner competitive parameter and, in this case, the unique stable equilibrium state can be recalled from any initial key.
Xiang-Wei He, Chung-Ping Kwong, Zongben Xu
IEEE Trans. Neural Networks3
1994 Asymmetric Bidirectional Associative Memories
abstract
Bidirectional associative memory (BAM) is a potentially promising model for heteroassociative memories. However, its applications are severely restricted to networks with logical symmetry of interconnections and pattern orthogonality or small pattern size. Although the restrictions on pattern orthogonality and pattern size can be relaxed to a certain extent, all previous efforts are at the cost of increase in connection complexity. In this paper, a new modification of the BAM is made and a new model named asymmetric bidirectional associative memory (ABAM) is proposed. This model not only can cater for the logical asymmetry of interconnections but also is capable of accommodating a larger number of non-orthogonal patterns. Furthermore, all these properties of the ABAM are achieved without increasing the connection complexity of the network. Theoretical analysis and simulation results all demonstrate that the ABAM indeed outperforms the BAM and its existing variants in all aspects of storage capacity, error-correcting capability and convergence.>
Zongben Xu, Yee Leung, Xiang-Wei He
IEEE Trans. Syst. Man Cybern. Syst.1