Yongyong Chen

dblp:196/7154 · DBLP profile ↗
← Back
113ranked-venue papers
20as first author
99since 2021 · last 2026
0000-0003-1970-1993ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 11 first-author · 51 since 2021Artificial intelligence and machine learning · 44 · 7 first-author · 42 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Security and privacy · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adversarial flow-based generative models for visible-to-Infrared person re-Identification
Honghu Pan, Yongyong Chen, Xin Li 0034, Zhenyu He 0001
Pattern Recognit.2
2026 Text-to-motion retrieval by text-to-motion generation
Honghu Pan, Yongyong Chen
Pattern Recognit.4
2026 Fine-grained tensor completion for incomplete multi-view clustering
Chong Peng 0001, Chundan Liu, Yongyong Chen, Zhao Kang 0001, Junyu Dong, Guiyuan Jiang, Chenglizhao Chen
Pattern Recognit.4
2026 Towards efficient and robust correntropy-based anchor tensor learning for multi-view subspace clustering
Shuqin Wang 0001, Yongli Wang 0004, Fang Qiu, Yongyong Chen, Yi-Gang Cen, Fanghui Zhang
Signal Process.4
2026 Dual Tensor Low-Rank Representation for Subspace Clustering
abstract
Benefiting from the powerful tensor techniques, the tensor low-rank representation has been proposed to construct sophisticated subspace clustering models. Existing tensor low-rank representation methods predominantly rely on a single low-rank prior to reconstruct the row space, which is instrumental in determining the subspace membership of samples by the row space information. However, this strategy neglects the column space and would lead to a subspace information loss. To address this issue, we propose a Dual Tensor Low-Rank Representation method (DTLRR), the first subspace clustering framework to theoretically recover both row and column subspaces simultaneously. Particularly, not simply formulating a dual self-representation model, we instead prove the recovery of both row and column spaces via a unified theoretical framework. Then, we impose low-rank constraints on the two corresponding affinity tensors to effectively capture high-order correlations. Meanwhile, we theoretically demonstrate the existence of compact dictionary tensors within the dual self-representation framework, which effectively eliminates the null spaces of the affinity tensors and significantly reduces computational complexity. Furthermore, an efficient Alternating Direction Method of Multipliers (ADMM) algorithm is designed to solve the proposed DTLRR model with guaranteed convergence. Extensive experiments validate the superior performance of the proposed DTLRR in data clustering, hyperspectral image denoising, and hyperspectral anomaly detection.
Qiangqiang Shen, Yin-Ping Zhao, Yongyong Chen, Yongsheng Liang 0001, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Anchor-Induced Serial Tensor Representation for Multi-View Clustering
abstract
Multi-view clustering (MVC) has emerged as a powerful approach for integrating diverse sources of information from complex datasets. Nevertheless, existing methods struggle to accurately capture the global correlations and high-order structures in the data, and employ anchor-based techniques within a single dimension, limiting their representation. To address these issues, we propose an Anchor-induced Serial Tensor Representation (ASTR) framework, which effectively harnesses serial tensor representation to capture comprehensive multi-view information while reducing approximation errors and enhancing clustering performance. Specifically, ASTR begins with projection learning to explore low-dimensional latent spaces in multi-view data. Then, we introduce multi-anchor learning, where multiple anchor configurations are generated within the latent spaces, yielding a set of corresponding bipartite graphs. Besides, we organize these bipartite graphs into a sequence of global tensors, forming the serial tensor representation that encapsulates high-order inter- and intra-view relationships. Furthermore, we introduce the Laplace function to achieve a more accurate tensor rank approximation, complemented by a thorough theoretical analysis. Finally, a one-step clustering process, guided by adaptive weights, directly fuses the learned graphs to produce the final clustering indicator matrix. Experimental results demonstrate that ASTR possesses superior clustering accuracy and comparable efficiency.
Zonglin Liu 0001, Zhiwei Zhong 0001, Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen
IEEE Trans. Circuits Syst. Video Technol.6
2026 Deep LoRA-Unfolding Networks for Image Restoration
abstract
Deep unfolding networks (DUNs), combining conventional iterative optimization algorithms and deep neural networks into a multi-stage framework, have achieved remarkable accomplishments in Image Restoration (IR), such as spectral imaging reconstruction, compressive sensing and super-resolution. It unfolds the iterative optimization steps into a stack of sequentially linked blocks. Each block consists of a Gradient Descent Module (GDM) and a Proximal Mapping Module (PMM) which is equivalent to a denoiser from a Bayesian perspective, operating on Gaussian noise with a known level. However, existing DUNs suffer from two critical limitations: 1) their PMMs share identical architectures and denoising objectives across stages, ignoring the need for stage-specific adaptation to varying noise levels; and 2) their chain of structurally repetitive blocks results in severe parameter redundancy and high memory consumption, hindering deployment in large-scale or resource-constrained scenarios. To address these challenges, we introduce generalized Deep Low-rank Adaptation (LoRA) Unfolding Networks for image restoration, named LoRun, harmonizing denoising objectives and adapting different denoising levels between stages with compressed memory usage for more efficient DUN. LoRun introduces a novel paradigm where a single pretrained base denoiser is shared across all stages, while lightweight, stage-specific LoRA adapters are injected into the PMMs to dynamically modulate denoising behavior according to the noise level at each unfolding step. This design decouples the core restoration capability from task-specific adaptation, enabling precise control over denoising intensity without duplicating full network parameters and achieving up to $N$ times parameter reduction for an $N$ -stage DUN with on-par or better performance. Extensive experiments conducted on three IR tasks validate the efficiency of our method.
Xiangming Wang, Haijin Zeng, Benteng Sun, Jiezhang Cao, Kai Zhang 0008, Qiangqiang Shen, Yongyong Chen
IEEE Trans. Image Process.7
2025 Global Graph Propagation with Hierarchical Information Transfer for Incomplete Contrastive Multi-view Clustering
abstract
Incomplete multi-view clustering has become one of the important research problems due to the extensive missing multi-view data in the real world. Although the existing methods have made great progress, there are still some problems: 1) most methods cannot effectively mine the information hidden in the missing data; 2) most methods typically divide representation learning and clustering into two separate stages, but this may affect the clustering performance as the clustering results directly depend on the learned representation. To address these problems, we propose a novel incomplete multi-view clustering method with hierarchical information transfer. Firstly, we design the view-specific Graph Convolutional Networks (GCN) to obtain the representation encoding the graph structure, which is then fused into the consensus representation. Secondly, considering that one layer of GCN transfers one-order neighbor node information, the global graph propagation with the consensus representation is proposed to handle the missing data and learn deep representation. Finally, we design a weight-sharing pseudo-classifier with contrastive learning to obtain an end-to-end framework that combines view-specific representation learning, global graph propagation with hierarchical information transfer, and contrastive clustering for joint optimization. Extensive experiments conducted on several commonly-used datasets demonstrate the effectiveness and superiority of our method in comparison with other state-of-the-art approaches.
Guoqing Chao, Kaixin Xu, Xijiong Xie, Yongyong Chen
AAAI4
2025 OTLRM: Orthogonal Learning-based Low-Rank Metric for Multi-Dimensional Inverse Problems
abstract
In real-world scenarios, complex data such as multispectral images and multi-frame videos inherently exhibit robust low-rank property. This property is vital for multi-dimensional inverse problems, such as tensor completion, spectral imaging reconstruction, and multispectral image denoising. Existing tensor singular value decomposition (t-SVD) definitions rely on hand-designed or pre-given transforms, which lack flexibility for defining tensor nuclear norm (TNN). The TNN-regularized optimization problem is solved by the singular value thresholding (SVT) operator, which leverages the t-SVD framework to obtain the low-rank tensor. However, it's quite complicated to introduce SVT into deep neural network due to the numerical instability problem in solving the derivatives of the eigenvectors. In this paper, we introduce a novel data-driven generative low-rank t-SVD model based on the learnable orthogonal transform, which can be naturally solved under its representation. Prompted by the linear algebra theorem of the Householder transformation, our learnable orthogonal transform is achieved by constructing an endogenously orthogonal matrix adaptable to neural networks, optimizing it as arbitrary orthogonal matrices. Additionally, we propose a low-rank solver as a generalization of SVT, which utilizes an efficient representation of generative networks to obtain low-rank structures. Extensive experiments highlight its significant restoration enhancements.
Xiangming Wang, Haijin Zeng, Jiaoyang Chen, Sheng Liu 0033, Yongyong Chen, Guoqing Chao
AAAI5
2025 Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks
abstract
Dynamic image degradations, including noise, blur and lighting inconsistencies, pose significant challenges in image restoration, often due to sensor limitations or adverse environmental conditions. Existing Deep Unfolding Networks (DUNs) offer stable restoration performance but require manual selection of degradation matrices for each degradation type, limiting their adaptability across diverse scenarios. To address this issue, we propose the Vision-Language-guided Unfolding Network (VLU-Net), a unified DUN framework for handling multiple degradation types simultaneously. VLU-Net leverages a VisionLanguage Model (VLM) refined on degraded image-text pairs to align image features with degradation descriptions, selecting the appropriate transform for target degradation. By integrating an automatic VLM-based gradient estimation strategy into the Proximal Gradient Descent (PGD) algorithm, VLU-Net effectively tackles complex multi-degradation restoration tasks while maintaining interpretability. Furthermore, we design a hierarchical feature unfolding structure to enhance VLU-Net framework, efficiently synthesizing degradation patterns across various levels. VLU-Net is the first all-in-one DUN framework and outperforms current leading one-by-one and all-in-one end- to-end methods by 3.74 dB on the SOTS dehazing dataset and 1.70 dB on the Rain100L deraining dataset.
Haijin Zeng, Xiangming Wang, Yongyong Chen, Jingyong Su
CVPR3
2025 Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing
abstract
Quad Bayer demosaicing is the central challenge for enabling the widespread application of Hybrid Event-based Vision Sensors (HybridEVS). Although existing learning-based methods that leverage long-range dependency modeling have achieved promising results, their complexity severely limits deployment on mobile devices for real-world applications. To address these limitations, we propose a lightweight Mamba-based binary neural network designed for efficient and high-performing demosaicing of HybridEVS RAW images. First, to effectively capture both global and local dependencies, we introduce a hybrid Binarized Mamba-Transformer architecture that combines the strengths of the Mamba and Swin Transformer architectures. Next, to significantly reduce computational complexity, we propose a binarized Mamba (Bi-Mamba), which binarizes all projections while retaining the core Selective Scan in full precision. Bi-Mamba also incorporates additional global visual information to enhance global context and mitigate precision loss. We conduct quantitative and qualitative experiments to demonstrate the effectiveness of BMTNet in both performance and computational efficiency, providing a lightweight demosaicing solution suited for real-world edge devices. Our codes and models are available at https://github.com/Clausy9/BMTNet.
Haijin Zeng, Yunfan Lu, Tong Shao, Yongyong Chen, Jingyong Su
CVPR6
2025 Time-Efficient Uncertainty Estimation Based on Target Networks in Deep Reinforcement Learning
abstract
Driven by the growing challenges of safety, exploration, generalization, and robustness in real-world applications, uncertainty estimation has emerged as a significant research domain in deep reinforcement learning (DRL). Uncertainty estimation theories can be broadly categorized into two main approaches: Bayesian theory-based methods and ensemble theory-based methods. However, Bayesian theory-based methods try to approximate the posterior distribution to estimate the uncertainties, and ensemble theory-based methods train multiple neural networks as a network ensemble to approximate the real distribution. They both cause a significant increase in training time. To address the above problems, we propose a target network-based approach for efficient uncertainty estimation (UEDQN), which uses time-updated target networks to replace the training of extra network. Specifically, UEDQN employs Bayesian theory to model the uncertainties, thereby distinguishing aleatoric uncertainty and epistemic uncertainty. Then, the two uncertainties are quantified and estimated using two periodically updated target networks. Finally, estimated uncertainties are used for exploration to accelerate the efficiency of training. The experimental results demonstrate the superior training efficiency of our method in environments with uncertainties and the excellent performance of the final trained agents.
Xionglue Li, Yongguang Wang, Yongyong Chen
ICIP4
2025 Spectral Compressive Imaging via Unmixing-driven Subspace Diffusion Refinement
abstract
Spectral Compressive Imaging (SCI) reconstruction is inherently ill-posed because a single observation admits multiple plausible reconstructions. Traditional deterministic methods struggle to effectively recover high-frequency details. Although diffusion models offer promising solutions to this challenge, their application is constrained by the limited training data and high computational demands associated with multispectral images (MSIs), making direct diffusion training impractical. To address these issues, we propose a novel Predict-and-unmixing-driven-Subspace-Refine framework (PSR-SCI). This framework begins with a light-weight predictor that produces an initial, rough estimate of the MSI. Subsequently, we introduce a unmixing-driven reversible spectral embedding module that decomposes the MSI into subspace images and spectral coefficients. This compact representation facilitates the adaptation of pre-trained RGB diffusion models and focuses refinement processes on high-frequency details, thereby enabling efficient diffusion generation with minimal MSI data. Additionally, we design a high-dimensional guidance mechanism enforcing SCI consistency during sampling. The refined subspace image is then reconstructed back into an MSI using the reversible embedding, yielding the final MSI with full spectral resolution. Experimental results on the standard KAIST and zero-shot datasets NTIRE, ICVL, and Harvard show that PSR-SCI enhances overall visual quality and delivers PSNR and SSIM results competitive with state-of-the-art diffusion, transformer, and deep-unfolding baselines. This framework provides a robust alternative to traditional deterministic SCI reconstruction methods. Code and models are available at [https://github.com/SMARK2022/PSR-SCI](https://github.com/SMARK2022/PSR-SCI).
Haijin Zeng, Benteng Sun, Yongyong Chen, Jingyong Su, Yong Xu 0001
ICLR3
2025 Multi-modal data augmentation based on masked modeling for image-text retrieval
Guoqing Chao, Yongyong Chen, Xijiong Xie
Knowl. Based Syst.4
2025 A lightweight dual-student mean teacher semi-supervised semantic segmentation method for skin lesions
Jindian Lu, Yongyong Chen, Yuanhaonan Deng, Binghui Zhao, Lixia Xue
Neural Networks3
2025 Adaptively robust high-order tensor factorization for low-rank tensor reconstruction
Yongyong Chen, Weihua Zhao
Pattern Recognit.2
2025 When non-local similarity meets tensor factorization: A patch-wise method for hyperspectral anomaly detection
Lixiang Meng, Yanhui Xu, Qiangqiang Shen, Yongyong Chen
Signal Process.4
2025 Reliable Entropy-Induced Anchor Learning for Incomplete Multi-View Subspace Clustering
abstract
Under large-scale data with missing views, fast incomplete multi-view clustering (IMVC) with anchor learning is of critical importance due to its linear complexity$\mathcal {O}(n)$. However, existing anchor-based methods only explore the column orthogonality of anchor points, where their arbitrary column orthogonal basis vectors have weak constraint relationships with real samples and significant deviations from more representative anchors, thereby impeding the precise representation of sample similarities. To solve this issue, we propose a Reliable Entropy-induced anchor learning for incomplete Multi-view subspace Clustering (REMC), which performs an entropy approximation term to learn more representative anchors, and we prove that the information entropy minimization can be relaxed into the$\ell _{2,1}$-norm paradigm. Specifically, the proposed REMC first integrates anchor learning and subspace clustering to produce multiple view-specific bipartite graphs and capture the high-order correlations by imposing these bipartite graphs with the tensor nuclear norm. Then, we fuse all the view-specific bipartite graphs to build a consensus bipartite graph with entropy approximation regularization, and hence the proposed REMC can produce a more discriminative similarity graph, preserving each non-zero element in its column close to 1, while the other elements are approaching 0. Besides, an efficient algorithm is designed to solve the proposed REMC. Numerous results show the superior performance of our method on both the complete and incomplete data.
Qiangqiang Shen, Zihou Guo, Yanhui Xu, Yongyong Chen, Shiqi Wang 0001, Yongsheng Liang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Sleep Stage Classification With Multi-Modal Fusion and Denoising Diffusion Model
abstract
Sleep stage classification plays a crucial role in sleep quality assessment and sleep disorder prevention. Nowadays, many studies have developed algorithms for this purpose, but they still face two challenges. The first is noise in physiological signals from various devices. The second challenge is that most studies simply concatenate multi-modal features without considering their correlations. To this end, we propose a framework, namely Diff-SleepNet, to efficiently classify sleep stages from multi-modal input. This framework begins with a diffusion model with peak signal-to-noise ratio (PNSR) loss function that adaptively filters noise. The filtered signals are then transformed into a multi-view spectrum through data pre-processing. These spectra are processed by a transformer-based backbone to extract multi-modal features. The production is fed into the following multi-scale attention module for robust feature fusion. The sleep stage category is finally determined by a fully connected layer. Our framework is trained and validated on three typical datasets, i.e., SHHS, Sleep-EDF-SC, and Sleep-EDF-X. Experimental results demonstrate that it is effective and has advantages over other peer methods.
Fengyu Cong, Yongyong Chen, Junxin Chen 0001
IEEE J. Biomed. Health Informatics3
2025 Smooth Tensor Qatar Riyal Decomposition for Dynamic MRI Reconstruction
abstract
Dynamic magnetic resonance imaging (dMRI) speed and imaging quality have always been a crucial issue in medical imaging research. Most existing methods characterize the tensor rank-based minimization to reconstruct dMRI from sampling $\bf k$-$t$ space data. However, (1) these approaches that unfold the tensor along each dimension destroy the inherent structure of dMR images. (2) they focus on preserving global information only, while ignoring the local details reconstruction such as the spatial piece-wise smoothness and sharp boundaries. To overcome these obstacles, we suggest a novel low-rank tensor decomposition approach by integrating tensor Qatar Riyal (QR) decomposition, low-rank tensor nuclear norm, and asymmetric total variation to reconstruct dMRI, named TQRTV. Specifically, while preserving the tensor inherent structure by utilizing tensor nuclear norm minimization to approximate tensor rank, QR decomposition reduces the dimensions in the low-rank constraint term, thereby improving the reconstruction performance. TQRTV further exploits the asymmetric total variation regularizer to capture local details. Numerical experiments demonstrate that the proposed reconstruction approach is superior to the existing ones.
Yongyong Chen, Haijin Zeng, Jingyong Su
IEEE J. Biomed. Health Informatics2
2024 Block Image Compressive Sensing with Local and Global Information Interaction
abstract
Block image compressive sensing methods, which divide a single image into small blocks for efficient sampling and reconstruction, have achieved significant success. However, these methods process each block locally and thus disregard the global communication among different blocks in the reconstruction step. Existing methods have attempted to address this issue with local filters or by directly reconstructing the entire image, but they have only achieved insufficient communication among adjacent pixels or bypassed the problem. To directly confront the communication problem among blocks and effectively resolve it, we propose a novel approach called Block Reconstruction with Blocks' Communication Network (BRBCN). BRBCN focuses on both local and global information, while further taking their interactions into account. Specifically, BRBCN comprises dual CNN and Transformer architectures, in which CNN is used to reconstruct each block for powerful local processing and Transformer is used to calculate the global communication among all the blocks. Moreover, we propose a global-to-local module (G2L) and a local-to-global module (L2G) to effectively integrate the representations of CNN and Transformer, with which our BRBCN network realizes the bidirectional interaction between local and global information. Extensive experiments show our BRBCN method outperforms existing state-of-the-art methods by a large margin. The code is available at https://github.com/kongxiuxiu/BRBCN
Xiaoyu Kong, Yongyong Chen, Feng Zheng 0001, Zhenyu He 0001
AAAI2
2024 Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset Selection
abstract
Training neural networks with good generalization requires large computational costs in many deep learning methods due to large-scale datasets and over-parameterized models. Despite the emergence of a number of coreset selection methods to reduce the computational costs, the problem of coreset distribution bias, i.e., the skewed distribution between the coreset and the entire dataset, has not been well studied. In this paper, we find that the closer the feature distribution of the coreset is to that of the entire dataset, the better the generalization performance of the coreset, particularly under extreme pruning. This motivates us to propose a simple yet effective method for coreset selection to alleviate the distribution bias between the coreset and the entire dataset, called feature distribution matching (FDMat). Unlike gradient-based methods, which selects samples with larger gradient values or approximates gradient values of the entire dataset, FDMat aims to select coreset that is closest to feature distribution of the entire dataset. Specifically, FDMat transfers coreset selection as an optimal transport problem from the coreset to the entire dataset in feature embedding spaces. Moreover, our method shows strong robustness due to the removal of samples far from the distribution, especially for the entire dataset containing noisy and class-imbalanced samples. Extensive experiments on multiple benchmarks show that FDMat can improve the performance of coreset selection than existing coreset methods. The code is available at https://github.com/successhaha/FDMat.
Weiwei Xiao, Yongyong Chen, Qiben Shan, Jingyong Su
AAAI2
2024 DiffSCI: Zero-Shot Snapshot Compressive Imaging via Iterative Spectral Diffusion Model
abstract
This paper endeavors to advance the precision of snap-shot compressive imaging (SCI) reconstruction for multi-spectral image (MSI). To achieve this, we integrate the ad-vantageous attributes of established SCI techniques and an image generative model, propose a novel structured zero-shot diffusion model, dubbed DiffSCI. DiffSCI leverages the structural insights from the deep prior and optimization-based methodologies, complemented by the generative ca-pabilities offered by the contemporary denoising diffusion model. Specifically, firstly, we employ a pre-trained diffusion model, which has been trained on a substantial corpus of RGB images, as the generative denoiser within the Plug-and-Play framework for the first time. This integration allows for the successful completion of SCI reconstruction, especially in the case that current methods struggle to address effectively. Secondly, we systematically account for spectral band correlations and introduce a robust methodology to mitigate wavelength mismatch, thus enabling seamless adaptation of the RGB diffusion model to MSIs. Thirdly, an accelerated algorithm is implemented to expedite the resolution of the data subproblem. This augmentation not only accelerates the convergence rate but also elevates the quality of the reconstruction process. We present extensive testing to show that DiffSCI exhibits discernible performance en-hancements over prevailing self-supervised and zero-shot approaches, surpassing even supervised transformer coun-terparts across both simulated and real datasets. Code is at https://github.com/PAN083/DiffSCI.
Zhenghao Pan, Haijin Zeng, Jiezhang Cao, Kai Zhang 0008, Yongyong Chen
CVPR5
2024 Fine-Grained Bipartite Concept Factorization for Clustering
abstract
In this paper, we propose a novel concept factorization method that seeks factor matrices using a cross-order positive semi-definite neighbor graph, which provides comprehensive and complementary neighbor information of the data. The factor matrices are learned with bipartite graph partitioning, which exploits explicit cluster structure of the data and is more geared towards clustering application. We develop an effective and efficient optimization algorithm for our method, and provide elegant theoretical results about the convergence. Extensive experimental results confirm the effectiveness of the proposed method.
Chong Peng 0001, Pengfei Zhang 0016, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
CVPR3
2024 Unmixing Diffusion for Self-Supervised Hyperspectral Image Denoising
abstract
Hyperspectral images (HSIs) have extensive applications in various fields such as medicine, agriculture, and industry. Nevertheless, acquiring high signal-to-noise ratio HSI poses a challenge due to narrow-band spectral filtering. Consequently, the importance of HSI denoising is substantial, especially for snapshot hyperspectral imaging technology. While most previous HSI denoising methods are supervised, creating supervised training datasets for the diverse scenes, hyperspectral cameras, and scan parameters is impractical. In this work, we present Diff-Unmix, a self-supervised denoising method for HSI using diffusion denoising generative models. Specifically, Diff-Unmix addresses the challenge of recovering noise-degraded HSI through a fusion of Spectral Unmixing and conditional abundance generation. Firstly, it employs a learnable block-based spectral unmixing strategy, complemented by a pure transformer-based backbone. Then, we introduce a self-supervised generative diffusion network to enhance abundance maps from the spectral unmixing block. This network reconstructs noise-free Unmixing probability distributions, effectively mitigating noise-induced degradations within these components. Finally, the reconstructed HSI is reconstructed through unmixing reconstruction by blending the diffusion-adjusted abundance map with the spectral endmembers. Experimental results on both simulated and real-world noisy datasets show that Diff-Unmix achieves state-of-the-art performance.
Haijin Zeng, Jiezhang Cao, Kai Zhang 0008, Yongyong Chen, Hiêp Quang Luong, Wilfried Philips
CVPR4
2024 Dual Prior Unfolding for Snapshot Compressive Imaging
abstract
Recently, deep unfolding methods have achieved remarkable success in the realm of Snapshot Compressive Imaging (SCI) reconstruction. However, the existing methods all follow the iterative framework of a single image prior, which limits the efficiency of the unfolding methods and makes it a problem to use other priors simply and effectively. To break out of the box, we derive an effective Dual Prior Unfolding (DPU), which achieves the joint utilization of multiple deep priors and greatly improves iteration efficiency. Our unfolding method is implemented through two parts, i.e., Dual Prior Framework (DPF) and Focused Attention (FA). In brief, in addition to the normal image prior, DPF introduces a residual into the iteration formula and constructs a degraded prior for the residual by considering various degradations to establish the unfolding framework. To improve the effectiveness of the image prior based on self-attention, FA adopts a novel mechanism inspired by PCA denoising to scale and filter attention, which lets the attention focus more on effective features with little computation cost. Besides, an asymmetric backbone is proposed to further improve the efficiency of hierarchical self-attention. Remarkably, our 5-stage DPU achieves state-of-the-art (SOTA) performance with the least FLOPs and parameters compared to previous methods, while our 9-stage DPU significantly outperforms other unfolding methods with less computational requirement. https: / /gi thub. com/ZhangJC-2k/DPU
Jiancheng Zhang 0003, Haijin Zeng, Jiezhang Cao, Yongyong Chen, Dengxiu Yu, Yin-Ping Zhao
CVPR4
2024 Improving Spectral Snapshot Reconstruction with Spectral-Spatial Rectification
abstract
How to effectively utilize the spectral and spatial char-acteristics of Hyperspectral Image (HSI) is always a key problem in spectral snapshot reconstruction. Recently, the spectra-wise transformer has shown great potential in capturing inter-spectra similarities of HSI, but the classic design of the transformer, i.e., multi-head division in the spectral (channel) dimension hinders the modeling of global spectral information and results in mean effect. In addition, previous methods adopt the normal spatial priors without taking imaging processes into account and fail to address the unique spatial degradation in snapshot spectral reconstruction. In this paper, we analyze the influence of multi-head division and propose a novel Spectral-Spatial Recti-fication (SSR) method to enhance the utilization of spectral information and improve spatial degradation. Specifically, SSR includes two core parts: Window-based Spectra-wise Self-Attention (WSSA) and spAtial Rectification Block (ARB). WSSA is proposed to capture global spectral in-formation and account for local differences, whereas ARB aims to mitigate the spatial degradation using a spatial alignment strategy. The experimental results on simulation and real scenes demonstrate the effectiveness of the proposed modules, and we also provide models at multiple scales to demonstrate the superiority of our approach. https://github.com/ZhangJC-2k/SSR
Jiancheng Zhang 0003, Haijin Zeng, Yongyong Chen, Dengxiu Yu, Yin-Ping Zhao
CVPR3
2024 SAH-SCI: Self-supervised Adapter for Efficient Hyperspectral Snapshot Compressive Imaging
Haijin Zeng, Yongyong Chen, Youfa Liu, Chong Peng 0001, Jingyong Su
ECCV (64)3
2024 Deep Unfolding 3D Non-Local Transformer Network for Hyperspectral Snapshot Compressive Imaging
abstract
Hyperspectral compressive imaging has shown remarkable advancements through the adoption of deep unfolding frameworks, which integrate the proximal mapping prior into the data fidelity term to formulate the reconstruction problem. However, existing technologies still face challenges in effectively capturing spatial-spectral features during the iterative deep prior learning stage, leading to unsatisfactory performance degradation. To address this issue, we propose a deep unfolding 3D non-local transformer (3DNLT) network for hyperspectral compressive imaging. A learnable half-quadratic splitting (HQS) algorithm is utilized to iteratively update the linear projection. Furthermore, a 3D non-local attention ushaped transformer is presented as the deep proximal mapping prior module to obtain the spatial-spectral long-range dependency features, leading to enhance the network’s ability to capture fine-grained hyperspectral and spatial details. Experimental results on both synthetic and real hyperspectral image reconstruction have demonstrated the superior performance of the 3DNLT network compared to state-of-the-art methods.
Yongyong Chen, Bingzhi Chen, Yicong Zhou
ICME3
2024 Cross-View Diversity Embedded Consensus Learning for Multi-View Clustering
Chong Peng 0001, Kai Zhang 0008, Yongyong Chen, Chenglizhao Chen, Qiang Shawn Cheng
IJCAI3
2024 MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging
abstract
Color video snapshot compressive imaging (SCI) employs computational imaging techniques to capture multiple sequential video frames in a single Bayer-patterned measurement. With the increasing popularity of quad-Bayer pattern in mainstream smartphone cameras for capturing high-resolution videos, mobile photography has become more accessible to a wider audience. However, existing color video SCI reconstruction algorithms are designed based on the traditional Bayer pattern. When applied to videos captured by quad-Bayer cameras, these algorithms often result in color distortion and ineffective demosaicing, rendering them impractical for primary equipment. To address this challenge, we propose the MambaSCI method, which leverages the Mamba and UNet architectures for efficient reconstruction of quad-Bayer patterned color video SCI. To the best of our knowledge, our work presents the first algorithm for quad-Bayer patterned SCI reconstruction, and also the initial application of the Mamba model to this task. Specifically, we customize Residual-Mamba-Blocks, which residually connect the Spatial-Temporal Mamba (STMamba), Edge-Detail-Reconstruction (EDR) module, and Channel Attention (CA) module. Respectively, STMamba is used to model long-range spatial-temporal dependencies with linear complexity, EDR is for better edge-detail reconstruction, and CA is used to compensate for the missing channel information interaction in Mamba model. Experiments demonstrate that MambaSCI surpasses state-of-the-art methods with lower computational and memory costs. PyTorch style pseudo-code for the core modules is provided in the supplementary materials. Code is at https://github.com/PAN083/MambaSCI.
Zhenghao Pan, Haijin Zeng, Jiezhang Cao, Yongyong Chen, Kai Zhang 0008, Yong Xu 0001
NeurIPS4
2024 Class-Agnostic Detection of Unknown Objects from Foreground Improves Robust Open World Object Detection
Yongyong Chen, Zimu Zheng, Jingyong Su
PRCV (12)3
2024 Convex-Concave Tensor Robust Principal Component Analysis
Youfa Liu, Bo Du 0001, Yongyong Chen, Lefei Zhang, Mingming Gong, Dacheng Tao
Int. J. Comput. Vis.3
2024 Bilevel fuzzy clustering via adaptive similarity graphs fusion
Yin-Ping Zhao, Xiangfeng Dai, Yongyong Chen, Chuanbin Zhang, Long Chen 0001
Inf. Sci.3
2024 Robust multiple subspaces transfer for heterogeneous domain adaptation
Youfa Liu, Bo Du 0001, Yongyong Chen, Lefei Zhang
Pattern Recognit.3
2024 Enhancing Inter-Class Separability With High-Order Strangers for Multi-View Clustering
abstract
Multi-view clustering has attracted extensive attention in recent years, which aims at integrating data from different views to improve the clustering performance. In this letter, we propose a novel approach for multi-view clustering. We propose to leverage high-order stranger information of the samples with the aid of Markov random walks to enhance inter-class separability of representation matrix in each view. Then, we seek a direct and intuitive clustering interpretation through view-specific spectral embeddings and cross-view spectral rotation fusion with auto-adjusted weights. Extensive experimental results confirm the effectiveness of our method.
Chundan Liu, Yongyong Chen, Junyu Dong, Chong Peng 0001
IEEE Signal Process. Lett.3
2024 A Robust Deep Learning Framework Based on Spectrograms for Heart Sound Classification
abstract
Heart sound analysis plays an important role in early detecting heart disease. However, manual detection requires doctors with extensive clinical experience, which increases uncertainty for the task, especially in medically underdeveloped areas. This paper proposes a robust neural network structure with an improved attention module for automatic classification of heart sound wave. In the preprocessing stage, noise removal with Butterworth bandpass filter is first adopted, and then heart sound recordings are converted into time-frequency spectrum by short-time Fourier transform (STFT). The model is driven by STFT spectrum. It automatically extracts features through four down sample blocks with different filters. Subsequently, an improved attention module based on Squeeze-and-Excitation module and coordinate attention module is developed for feature fusion. Finally, the neural network will give a category for heart sound waves based on the learned features. The global average pooling layer is adopted for reducing the model's weight and avoiding overfitting, while focal loss is further introduced as the loss function to minimize the data imbalance problem. Validation experiments have been conducted on two publicly available datasets, and the results well demonstrate the effectiveness and advantages of our method.
Junxin Chen 0001, Zhihuan Guo, Li-bo Zhang 0004, Yongyong Chen, Marcin Wozniak, Wei Wang 0077
IEEE Trans. Comput. Biol. Bioinform.6
2024 Self-Completed Bipartite Graph Learning for Fast Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC), excavating diversity and consistency from multiple incomplete views, has aroused widespread research enthusiasm. Nevertheless, most existing methods still encounter the following issues: 1) they generally concentrate on pair-wise instance correlation, which consumes at least a quadratic complexity and precludes them from applying at large scales; 2) they only concentrate on pair-wise instance relevance, whereas ignoring the discriminative correlation hidden across views. To overcome these drawbacks, we propose the Self-Completed Bipartite Graph Learning (SCBGL) method for fast IMVC, which adaptively learns a self-completed consensus bipartite graph with the guidance of global information. Specifically, SCBGL learns the consensus anchor matrix shared among diverse views and further constructs a consensus intra-view bipartite graph with missing instances to explore the diversity and complementarity underlying different views. Meanwhile, we concatenate all the multiple features with projection learning to learn global anchors that would be employed to construct an inter-view bipartite graph. Furthermore, SCBGL dexterously utilizes the abundant inter-view information to tutor the self-completion of the consensus intra-view bipartite graph. By devising an alternatively iterative strategy, we present an efficient algorithm, which enjoys a linear time complexity, to solve the proposed SCBGL model. Numerous experiments conducted on large-scale datasets substantiate the superior performance of the SCBGL beyond the state-of-the-arts.
Xiaojia Zhao, Qiangqiang Shen, Yongyong Chen, Yongsheng Liang 0001, Junxin Chen 0001, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.3
2024 Partial Tubal Nuclear Norm-Regularized Multiview Subspace Learning
abstract
In this article, a unified multiview subspace learning model, called partial tubal nuclear norm-regularized multiview subspace learning (PTN2MSL), was proposed for unsupervised multiview subspace clustering (MVSC), semisupervised MVSC, and multiview dimension reduction. Unlike most of the existing methods which treat the above three related tasks independently, PTN2MSL integrates the projection learning and the low-rank tensor representation to promote each other and mine their underlying correlations. Moreover, instead of minimizing the tensor nuclear norm which treats all singular values equally and neglects their differences, PTN2MSL develops the partial tubal nuclear norm (PTNN) as a better alternative solution by minimizing the partial sum of tubal singular values. The PTN2MSL method was applied to the above three multiview subspace learning tasks. It demonstrated that these tasks organically benefited from each other and PTN2MSL has achieved better performance in comparison to state-of-the-art methods.
Yongyong Chen, Yin-Ping Zhao, Shuqin Wang 0001, Junxin Chen 0001, Zheng Zhang 0006
IEEE Trans. Cybern.1
2024 Double Discrete Cosine Transform-Oriented Multi-View Subspace Clustering
abstract
Low-rank tensor representation with the tensor nuclear norm has been rising in popularity in multi-view subspace clustering (MVSC), in which the tensor nuclear norm is commonly implemented using discrete Fourier transform (DFT). Unfortunately, existing DFT-oriented MVSC methods may provide unsatisfactory results since (1) DFT exploits complex arithmetic in the Fourier domain, usually resulting in high tubal tensor rank, and (2) local structural information is rarely considered. To solve these problems, in this paper, we propose a novel double discrete cosine transform (DCT)-oriented multi-view subspace clustering (D2CTMSC) method, in which the first DCT aims to derive the tensor nuclear norm without complex arithmetic while the second DCT aims to explore the local structure of the self-representation tensor, such that the essential low-rankness and sparsity embedding in multi-view features can be thoroughly exploited. Moreover, we design an effective alternating iteration strategy to solve the proposed model. Experimental results on four types of multi-view datasets (News stories, Face images, Scene images, and Generic objects) demonstrate the superiority of the D2CTMSC method compared with DFT-based methods and other state-of-the-art clustering methods.
Yongyong Chen, Shuqin Wang 0001, Yin-Ping Zhao, C. L. Philip Chen
IEEE Trans. Image Process.1
2024 Fine-Grained Essential Tensor Learning for Robust Multi-View Spectral Clustering
abstract
Multi-view subspace clustering (MVSC) has drawn significant attention in recent study. In this paper, we propose a novel approach to MVSC. First, the new method is capable of preserving high-order neighbor information of the data, which provides essential and complicated underlying relationships of the data that is not straightforwardly preserved by the first-order neighbors. Second, we design log-based nonconvex approximations to both tensor rank and tensor sparsity, which are effective and more accurate than the convex approximations. For the associated shrinkage problems, we provide elegant theoretical results for the closed-form solutions, for which the convergence is guaranteed by theoretical analysis. Moreover, the new approximations have some interesting properties of shrinkage effects, which are guaranteed by elegant theoretical results. Extensive experimental results confirm the effectiveness of the proposed method.
Chong Peng 0001, Kehan Kang, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
IEEE Trans. Image Process.3
2024 Pick-and-Place Transform Learning for Fast Multi-View Clustering
abstract
To manipulate large-scale data, anchor-based multi-view clustering methods have grown in popularity owing to their linear complexity in terms of the number of samples. However, these existing approaches pay less attention to two aspects. 1) They target at learning a shared affinity matrix by using the local information from every single view, yet ignoring the global information from all views, which may weaken the ability to capture complementary information. 2) They do not consider the removal of feature redundancy, which may affect the ability to depict the real sample relationships. To this end, we propose a novel fast multi-view clustering method via pick-and-place transform learning named PPTL, which could capture insightful global features to characterize the sample relationships quickly. Specifically, PPTL first concatenates all the views along the feature direction to produce a global matrix. Considering the redundancy of the global matrix, we design a pick-and-place transform with ℓ2,p-norm regularization to abandon the poor features and consequently construct a compact global representation matrix. Thus, by conducting anchor-based subspace clustering on the compact global representation matrix, PPTL can learn a consensus skinny affinity matrix with a discriminative clustering structure. Numerous experiments performed on small-scale to large-scale datasets demonstrate that our method is not only faster but also achieves superior clustering performance over state-of-the-art methods across a majority of the datasets.
Qiangqiang Shen, Yongyong Chen, Changqing Zhang 0002, Yonghong Tian 0001, Yongsheng Liang 0001
IEEE Trans. Image Process.2
2024 RCUMP: Residual Completion Unrolling With Mixed Priors for Snapshot Compressive Imaging
abstract
Deep unrolling-based snapshot compressive imaging (SCI) methods, which employ iterative formulas to construct interpretable iterative frameworks and embedded learnable modules, have achieved remarkable success in reconstructing 3-dimensional (3D) hyperspectral images (HSIs) from 2D measurement induced by coded aperture snapshot spectral imaging (CASSI). However, the existing deep unrolling-based methods are limited by the residuals associated with Taylor approximations and the poor representation ability of single hand-craft priors. To address these issues, we propose a novel HSI construction method named residual completion unrolling with mixed priors (RCUMP). RCUMP exploits a residual completion branch to solve the residual problem and incorporates mixed priors composed of a novel deep sparse prior and mask prior to enhance the representation ability. Our proposed CNN-based model can significantly reduce memory cost, which is an obvious improvement over previous CNN methods, and achieves better performance compared with the state-of-the-art transformer and RNN methods. In this work, our method is compared with the 9 most recent baselines on 10 scenes. The results show that our method consistently outperforms all the other methods while decreasing memory consumption by up to 80%.
Yin-Ping Zhao, Jiancheng Zhang 0003, Yongyong Chen, Zhen Wang 0004, Xuelong Li 0001
IEEE Trans. Image Process.3
2024 Data Completion-Guided Unified Graph Learning for Incomplete Multi-View Clustering
abstract
Due to its heterogeneous property, multi-view data has been widely concerned over single-view data for performance improvement. Unfortunately, some instances may be with partially available information because of some uncontrollable factors, for which the incomplete multi-view clustering (IMVC) problem is raised. IMVC aims to partition unlabeled incomplete multi-view data into their clusters by exploiting the heterogeneity of multi-view data and overcoming the difficulty of data loss. However, most existing IMVC methods like BSV, MIC, OMVC, and IVC tend to conduct basic completion processing on the input data, without taking advantage of the correlation between samples and information redundancy. To overcome the above issue, we propose one novel IMVC method named data completion-guided unified graph learning (DCUGL), which could complete the data of missing views and fuse multiple learned view-specific similarity matrices into one unified graph. Specifically, we first reduce the dimension of the input data to learn multiple view-specific similarity matrices. By stacking all view-specific similarity matrices, DCUGL constructs a third-order tensor with the low-rank constraint, such that sample correlation within and between views can be well explored. Finally, by dividing the original data into observed data and unobserved data, DCUGL can infer and complete the missing data according to the view-specific similarity matrices, and obtain a unified graph, which can be directly used for clustering. To solve the proposed model, we design an iterative algorithm, which is based on the alternating direction method of multipliers framework. The proposed model proves to be superior by benchmarking on six challenging datasets compared with state-of-the-art IMVC methods.
Tianhai Liang, Qiangqiang Shen, Shuqin Wang 0001, Yongyong Chen, Junxin Chen 0001
ACM Trans. Knowl. Discov. Data4
2024 When Channel Correlation Meets Sparse Prior: Keeping Interpretability in Image Compressive Sensing
abstract
Image compressive sensing (CS), recovering an unknown image by resorting to a small number of its measurements, has become an increasingly popular topic in multimedia technology and applications. For a better reconstruction, diverse priors, from the original sparse prior to the new deep prior, have been exploited. Despite the powerful learning capability and satisfactory reconstruction performance, the deep prior is known as a black box and loses clear interpretability. In this article, we first revisit image CS with different priors and observe that the method with hand-crafted sparse prior could still outperform state-of-art methods with deep prior or no prior when under the same settings, while the interpretability is well preserved. Then, towards a better performance of the sparse-prior-based method, we propose a Channel Adaptive Thresholding Network, namely CAT-Net. CAT-Net draws the support from channel correlation calculation to extend the single thresholding in the iterative soft thresholding algorithm (ISTA) into channel-wise thresholding. The channel adaptive thresholding conducts soft thresholding operation in each channel of the image features and can be adjusted adaptively to the inputs, which can reconstruct more precisely than a single static thresholding. The careful CAT operation can preserve patterns both in detail and holistically well. Experimental results demonstrate the proposed method outperforms the state-of-the-art image CS methods with both traditional and deep priors.
Xiaoyu Kong, Yongyong Chen, Zhenyu He 0001
IEEE Trans. Multim.2
2024 Robust Tensor Recovery for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering is gaining increased attention owing to its great success in mining underlying information from the missing views. However, the existing approaches still encounter two issues: 1) They generally do not give sufficient consideration to the robustness of incomplete multi-view data with noise; 2) They only exploit the low-rank structures in the intra-view graphs, while the low-rank priors embedded in inter-view graphs are ignored. To this end, we propose a Robust Tensor Recovery for Incomplete Multi-view Clustering (RIMC) method, which transforms the view-missing problem into the tensor graph recovery problem by manipulating the comprehensive low-rank priors. Specifically, RIMC first employs a marginalized denoising operation to construct robust graphs and further builds a tensor graph by stacking these robust graphs. Then, we develop a novel tensor completion to recover the tensor graph by performing comprehensive low-rank priors: low-rank structures in the inter-view graphs (i.e., horizontal and lateral slices); low-rank structures in the intra-view graphs (i.e., frontal slices). Meanwhile, we integrate the tensor completion and spectral clustering to learn a unified indicator matrix. Extensive experiments show the promising performance of our method.
Qiangqiang Shen, Yongsheng Liang 0001, Yongyong Chen, Zhenyu He 0001
IEEE Trans. Multim.4
2024 Tensor Learning Meets Dynamic Anchor Learning: From Complete to Incomplete Multiview Clustering
abstract
Multiview clustering (MVC), which can dexterously uncover the underlying intrinsic clustering structures of the data, has been particularly attractive in recent years. However, previous methods are designed for either complete or incomplete multiview only, without a unified framework that handles both tasks simultaneously. To address this issue, we propose a unified framework to efficiently tackle both tasks in approximately linear complexity, which integrates tensor learning to explore the inter-view low-rankness and dynamic anchor learning to explore the intra-view low-rankness for scalable clustering (TDASC). Specifically, TDASC efficiently learns smaller view-specific graphs by anchor learning, which not only explores the diversity embedded in multiview data, but also yields approximately linear complexity. Meanwhile, unlike most current approaches that only focus on pair-wise relationships, the proposed TDASC incorporates multiple graphs into an inter-view low-rank tensor, which elegantly models the high-order correlations across views and further guides the anchor learning. Extensive experiments on both complete and incomplete multiview datasets clearly demonstrate the effectiveness and efficiency of TDASC compared with several state-of-the-art techniques.
Yongyong Chen, Xiaojia Zhao, Zheng Zhang 0006, Youfa Liu, Jingyong Su, Yicong Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2024 Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise Perspective
abstract
Arbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2.
Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.6
2024 Tensor Completion Using Bilayer Multimode Low-Rank Prior and Total Variation
abstract
In this article, we propose a novel bilayer low-rankness measure and two models based on it to recover a low-rank (LR) tensor. The global low rankness of underlying tensor is first encoded by LR matrix factorizations (MFs) to the all-mode matricizations, which can exploit multiorientational spectral low rankness. Presumably, the factor matrices of all-mode decomposition are LR, since local low-rankness property exists in within-mode correlation. In the decomposed subspace, to describe the refined local LR structures of factor/subspace, a new low-rankness insight of subspace: a double nuclear norm scheme is designed to explore the so-called second-layer low rankness. By simultaneously representing the bilayer low rankness of the all modes of the underlying tensor, the proposed methods aim to model multiorientational correlations for arbitrary N -way ( N ≥ 3 ) tensors. A block successive upper-bound minimization (BSUM) algorithm is designed to solve the optimization problem. Subsequence convergence of our algorithms can be established, and the iterates generated by our algorithms converge to the coordinatewise minimizers in some mild conditions. Experiments on several types of public datasets show that our algorithm can recover a variety of LR tensors from significantly fewer samples than its counterparts.
Haijin Zeng, Shaoguang Huang, Yongyong Chen, Sheng Liu 0033, Hiêp Quang Luong, Wilfried Philips
IEEE Trans. Neural Networks Learn. Syst.3
2024 Double High-Order Correlation Preserved Robust Multi-View Ensemble Clustering
abstract
Ensemble clustering (EC), utilizing multiple basic partitions (BPs) to yield a robust consensus clustering, has shown promising clustering performance. Nevertheless, most current algorithms suffer from two challenging hurdles: (1) a surge of EC-based methods only focus on pair-wise sample correlation while fully ignoring the high-order correlations of diverse views. (2) they deal directly with the co-association (CA) matrices generated from BPs, which are inevitably corrupted by noise and thus degrade the clustering performance. To address these issues, we propose a novel Double High-Order Correlation Preserved Robust Multi-View Ensemble Clustering (DC-RMEC) method, which preserves the high-order inter-view correlation and the high-order correlation of original data simultaneously. Specifically, DC-RMEC constructs a hypergraph from BPs to fuse high-level complementary information from different algorithms and incorporates multiple CA-based representations into a low-rank tensor to discover the high-order relevance underlying CA matrices, such that double high-order correlation of multi-view features could be dexterously uncovered. Moreover, a marginalized denoiser is invoked to gain robust view-specific CA matrices. Furthermore, we develop a unified framework to jointly optimize the representation tensor and the result matrix. An effective iterative optimization algorithm is designed to optimize our DC-RMEC model by resorting to the alternating direction method of multipliers. Extensive experiments on seven real-world multi-view datasets have demonstrated the superiority of DC-RMEC compared with several state-of-the-art multi-view ensemble clustering methods.
Xiaojia Zhao, Qiangqiang Shen, Youfa Liu, Yongyong Chen, Jingyong Su
ACM Trans. Multim. Comput. Commun. Appl.5
2023 PnP-AE: A Plug-and-Play Module for Volumetric Medical Image Segmentation
abstract
In recent years, 3D volumetric medical images have been widely used in clinical diagnosis, however, the popular 2D networks were reported unsuitable for segmenting them. In this direction, we propose a plug-and-play (PnP-AE) module to improve the performance of using 2D network for 3D medical image segmentation. Our method takes advantage of the intrinsic correlation between adjacent slices, by multiple encoders and fusion components to decouple plane feature extraction and depth information integration. In addition, the proposed weight sharing and feature storage strategies make PnP-AE extremely efficient. Our method is able to conveniently incorporate with mainstream 2D networks to segment 3D volumetric medical images. Experimental results demonstrate the excellent performance of our method. The source code is available at https://github.com/qklee-lz/PnP-AE.
Qiankun Li 0004, Xiaolong Huang 0001, Bo Fang 0005, Yongyong Chen, Junxin Chen 0001
BIBM5
2023 End-to-end XY Separation for Single Image Blind Deblurring
abstract
Single image blind deblurring, only exploiting a blurry observation to reconstruct the sharp image, is a popular yet challenging low-level vision task. Current state-of-the-art deblurring networks mainly follow the coarse-to-fine strategy for architecture design and utilize U-net or its variant, XYDeblur, as the basic units. However, the one-encoder-one-decoder and the recently proposed one-encoder-two-decoder structures of basic units both fail to comprehensively take advantage of the directional separability of 2D deblurring, which increases the learning content of networks, thus leading to performance degradation. To thoroughly decouple the deblurring into two spatially orthogonal parts, we propose a novel substitution for U-net and its variant, called XYU-net. Specifically, it consists of two structurally identical U-nets, named XU-net and YU-net. They share orthogonal parameters by rotating kernels and focus on restoring a 2D blurry image in two spatially orthogonal directions respectively, which not only brings efficiency enhancement but also maintains parameter number. To further reduce the graphics memory demand of XYU-net, we transfer some non-linear transform modules (NLTM) from the outside of the network to its inside and propose the modified version, called MXYU-net. Experimental results on three large blurry image datasets demonstrate the efficiency of XYU-net and MXYU-net compared with U-net and XYDeblur, both as standalone models and as basic units of advanced U-net-based deblurring networks.
Liuhan Chen, Yirou Wang, Yongyong Chen
ACM Multimedia3
2023 Laplacian regularized deep low-rank subspace clustering network
Yongyong Chen, Zhongyun Hua
Appl. Intell.1
2023 Global and local similarity learning in multi-kernel space for nonnegative matrix factorization
Chong Peng 0001, Xingrong Hou, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
Knowl. Based Syst.3
2023 HMM-GDAN: Hybrid multi-view and multi-scale graph duplex-attention networks for drug response prediction in cancer
Youfa Liu, Shufan Tong, Yongyong Chen
Neural Networks3
2023 Multi-granularity graph pooling for video-based person re-identification
Honghu Pan, Yongyong Chen, Zhenyu He 0001
Neural Networks2
2023 Saliency Transfer Learning and Central-Cropping Network for Prostate Cancer Classification
Mengpei Jia, Jihao Luo, Aijun Zhang, Yongyong Chen, Peipei Shan, Binghui Zhao
Neural Process. Lett.6
2023 Asymmetry total variation and framelet regularized nonconvex low-rank tensor completion
Yongyong Chen, Xiaojia Zhao, Haijin Zeng, Yanhui Xu, Junxing Chen
Signal Process.1
2023 Cross-Modality LGE-CMR Segmentation Using Image-to-Image Translation Based Data Augmentation
abstract
Accurate segmentation of ventricle and myocardium from the late gadolinium enhancement (LGE) cardiac magnetic resonance (CMR) is an important tool for myocardial infarction (MI) analysis. However, the complex enhancement pattern of LGE-CMR and the lack of labeled samples make its automatic segmentation difficult to be implemented. In this paper, we propose an unsupervised LGE-CMR segmentation algorithm by using multiple style transfer networks for data augmentation. It adopts two different style transfer networks to perform style transfer of the easily available annotated balanced-Steady State Free Precession (bSSFP)-CMR images. Then, multiple sets of synthetic LGE-CMR images are generated by the style transfer networks and used as the training data for the improved U-Net. The entire implementation of the algorithm does not require the labeled LGE-CMR. Validation experiments demonstrate the effectiveness and advantages of the proposed algorithm.
Wei Wang 0077, Xinhua Yu, Bo Fang 0005, Yongyong Chen, Wei Wei 0006, Junxin Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Enabling Large-Capacity Reversible Data Hiding Over Encrypted JPEG Bitstreams
abstract
Cloud computing offers advantages in handling the exponential growth of images but also entails privacy concerns on outsourced private images. Reversible data hiding (RDH) over encrypted images has emerged as an effective technique for securely storing and managing confidential images in the cloud. Most existing schemes only work on uncompressed images. However, almost all images are transmitted and stored in compressed formats such as JPEG. Recently, some RDH schemes over encrypted JPEG bitstreams have been developed, but these works have some disadvantages such as a small embedding capacity (particularly for low quality factors), damage to the JPEG format, and file size expansion. In this study, we propose a permutation-based embedding technique that allows the embedding of significantly more data than existing techniques. Using the proposed embedding technique, we further design a large-capacity RDH scheme over encrypted JPEG bitstreams, in which a grouping method is designed to boost the number of embeddable blocks. The designed RDH scheme allows a content owner to encrypt a JPEG bitstream before uploading it to a cloud server. The cloud server can embed additional data (e.g., copyright and identification information) into the encrypted JPEG bitstream for storage, management, or other processing purpose. A receiver can losslessly recover the original JPEG bitstream using a decryption key. Comprehensive evaluation results demonstrate that our proposed design can achieve approximately twice the average embedding capacity compared to the best prior scheme while preserving the file format without file size expansion.
Zhongyun Hua, Yifeng Zheng 0001, Yongyong Chen, Yuanman Li
IEEE Trans. Circuits Syst. Video Technol.4
2023 Pose-Aided Video-Based Person Re-Identification via Recurrent Graph Convolutional Network
abstract
Existing methods for video-based person re- identification (ReID) mainly learn the appearance feature of a given pedestrian via a feature extractor and a feature aggregator. However, the appearance models would fail to learn a large inter-class variance when different pedestrians have similar appearances. Considering that different pedestrians have different walking postures and body proportions, we propose to learn the discriminative pose feature beyond the appearance feature for video retrieval. Specifically, we implement a two-branch architecture to separately learn the appearance feature and pose feature, and then concatenate them together for inference. To learn the pose feature, we first detect the pedestrian pose in each frame through an off-the-shelf pose detector, and construct a temporal graph using the pose sequence. We then exploit a recurrent graph convolutional network (RGCN) to learn the node embeddings of the temporal pose graph, which devises a global information propagation mechanism to simultaneously achieve the neighborhood aggregation of intra-frame nodes and message passing among inter-frame graphs. Finally, we propose a dual-attention method (DAM) consisting of node-attention and time-attention to obtain the temporal graph representation from the node embeddings, where the self-attention mechanism is employed to learn the importance of each node and each frame. We verify the proposed method on three video-based ReID datasets, i.e., Mars, DukeMTMC and iLIDS-VID, whose experimental results demonstrate that the learned pose feature can effectively improve the performance of existing appearance models.
Honghu Pan, Qiao Liu 0001, Yongyong Chen, Yunqi He, Yuan Zheng 0002, Feng Zheng 0001, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Robustness Meets Low-Rankness: Unified Entropy and Tensor Learning for Multi-View Subspace Clustering
abstract
In this paper, we develop the weighted error entropy-regularized tensor learning method for multi-view subspace clustering (WETMSC), which integrates the noise disturbance removal and subspace structure discovery into one unified framework. Unlike most existing methods which focus only on the affinity matrix learning for the subspace discovery by different optimization models and simply assume that the noise is independent and identically distributed (i.i.d.), our WETMSC method adopts the weighted error entropy to characterize the underlying noise by assuming that noise is independent and piecewise identically distributed (i.p.i.d.). Meanwhile, WETMSC constructs the self-representation tensor by storing all self-representation matrices from the view dimension, preserving high-order correlation of views based on the tensor nuclear norm. To solve the proposed nonconvex optimization method, we design a half-quadratic (HQ) additive optimization technology and iteratively solve all subproblems under the alternating direction method of multipliers framework. Extensive comparison studies with state-of-the-art clustering methods on real-world datasets and synthetic noisy datasets demonstrate the ascendancy of the proposed WETMSC method.
Shuqin Wang 0001, Yongyong Chen, Zhiping Lin 0001, Yi-Gang Cen, Qi Cao 0002
IEEE Trans. Circuits Syst. Video Technol.2
2023 Deep and Low-Rank Quaternion Priors for Color Image Processing
abstract
Due to the physical nature of color images, color image processing such as denoising and inpainting has shown extensive and versatile possibilities over grayscale image processing. The monochromatic and the concatenation model have been widely used to process color images by processing each color channel independently or concatenating three color channels as one unified one and then used existing grayscale image processing methods directly without specific operations. These above schemes, however, have some limitations: (1) they would destroy the inherent correlation among three color channels since they cannot represent color images holistically; (2) they usually focus on one specific handcrafted prior such as smoothness, low-rankness, or even deep prior and thus failing to fuse deep and handcrafted priors of color images flexibly. To conquer these limitations, we propose one unified model to integrate deep prior and low-rank quaternion prior (DLRQP) for color image processing under the plug-and-play (PnP) framework. Specifically, the quaternion representation with low-rank constraint is introduced to denote the color image in a holistic way and one advanced denoiser is adopted to explore the deep prior in an iterative process. To tightly approximate the quaternion rank, one nonconvex penalty function is further utilized. We derive an alternate iterative approach to tackle the proposed model. We empirically demonstrate that our model can achieve superior performance over existing methods on both color image denoising and inpainting tasks.
Xiaoyu Kong, Qiangqiang Shen, Yongyong Chen, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2023 Deep Dynamic Memory Augmented Attentional Dictionary Learning for Image Denoising
abstract
Motivated by the advance of deep learning methods, deep unfolding methods such as deep convolutional dictionary learning have achieved great success in image denoising tasks. The main advantages are inheriting both the merits of deep learning (strong learning capacity) and traditional machine learning (powerful interpretable capacity). We observe that the update of dictionaries and coefficients is highly correlated with the previous iterative stage information for deep unfolding-based methods. However, most existing deep convolutional dictionary learning methods deal with each iteration step individually, ignoring the inner-memory within the stage and cross-memory across the stages. To alleviate these issues, we propose a dynamic inner-cross memory augmented attentional dictionary learning (M2ADL) network with attention guided residual connection module, which utilizes the previous important stage features such that better uncovering the inner-cross information. Specifically, the proposed inner-cross memory fully utilizes the previous stage’s hidden and last-layer information to learn the dictionary. In addition, we develop a dual attention-guided residual connection module to well exploit the deep feature learning ability to capture the spatial-spectral attention across the deep tensor-based features. Considerable experiments on both synthetic and real image datasets demonstrate the superiority of the proposed method over other state-of-the-art methods.
Yongyong Chen, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.2
2023 Matrix-Based Secret Sharing for Reversible Data Hiding in Encrypted Images
abstract
Traditional schemes for reversible data hiding in encrypted images (RDH-EI) focus on one data hider and cannot resist the single point of failure. Besides, the image security is determined by one party, rather than multiple parties. Thus, it is valuable to design RDH-EI schemes with multiple data hiders for stronger security. In this article, we propose a multiple data hiders-based RDH-EI scheme using a new secret sharing technique. First, we devise an$(r,n)$-threshold$(r\leq n)$matrix-based secret sharing (MSS) using matrix theory, and theoretically verify its efficacy and security properties. Then, using the MSS, we propose an$(r,n)$-threshold RDH-EI scheme called MSS-RDHEI. The content owner encrypts an image to be$n$encrypted images using the MSS with an encryption key, and outsources these encrypted images to$n$data hiders. Each data hider can embed some data, e.g., copyright and identification information, into the encrypted image for the purposes of storage, management, or other processing, and these data can also be losslessly extracted. An authorized receiver can recover the confidential image from$r$encrypted images. By designing, our MSS-RDHEI scheme can withstand$n-r$points of failure. Experimental results show that it ensures the image content confidentiality and achieves a much larger embedding capacity than state-of-the-art schemes.
Zhongyun Hua, Yifeng Zheng 0001, Yongyong Chen, Xinpeng Zhang 0001
IEEE Trans. Dependable Secur. Comput.6
2023 Generalized Nonconvex Low-Rank Tensor Representation for Hyperspectral Anomaly Detection
abstract
Low-rank tensor representation (LRTR) methods have attracted great interest for their powerful ability to separate backgrounds and anomalies. However, most of the current LRTR models use the popular and convex surrogate tensor nuclear norm to solve optimization problems, which results in a loose approximation and suboptimal solver for the original problem. Besides, most existing methods solve the nonconvex optimization problems case-by-case, consequently losing one unified solver. To solve the above issues, we propose the Generalized Nonconvex Low-rank Tensor Representation (GNLTR) for hyperspectral anomaly detection (HAD), a unified solver not case-by-case one of existing nonconvex optimization problems. Compared to the tensor nuclear norm, GNLTR contains many popular nonconvex penalty functions as tighter regularizers of the tensor tubal rank to constrain the low rank of the background. Moreover, theL2,1norm has been integrated into the GNLTR model for the sparse anomalies. For the optimization problem, it is handled quickly and efficiently through a well-organized alternating direction method of multipliers (ADMM). The experiments on several real-world hyperspectral data sets demonstrate the superior performance of the GNLTR model in comparison with some state-of-the-art anomaly detection models.
Qiangqiang Shen, Haijin Zeng, Yongyong Chen, Guangming Lu 0002
IEEE Trans. Geosci. Remote. Sens.4
2023 Cross-Scale-Guided Fusion Transformer for Disaster Assessment Using Satellite Imagery
abstract
When a disaster strikes, accurate disaster information and effective response are critical for saving lives and properties. High-resolution satellite (HRS) imagery provides valuable geographical information that can assist experts in analyzing damage levels in different areas and enacting appropriate relief plans. However, analyzing large HRS images is both time-consuming and inefficient, requiring efficient automated methods to replace expert analysis. Fortunately, deep learning methods have achieved impressive performance on HRS image processing tasks, considerably increasing automation levels. Despite this progress, most HRS-based damage assessment methods only consider a single time series of post-disaster images or simply integrate pre- and post-disaster images, lacking the integration of effective information between pre- and post-disaster images. To alleviate this problem, we propose a two-stage multi-scale fusion network that fully exploits the information contained in pre- and post-disaster images. Specifically, we employ a hierarchical Transformer to accurately locate buildings by pre-disaster images in the first stage, and then propose the guided fusion and cross-scale guided fusion modules in the second stage to efficiently utilize both pre- and post-disaster images. Our method outperforms state-of-the-art methods in building segmentation and building damage assessment on the xBD dataset, and exhibits improved generalization across diverse geographic regions and disaster types.
Weiwei Xiao, Jingyong Su, Yongyong Chen, Guofeng Cao
IEEE Trans. Geosci. Remote. Sens.3
2023 Hyperspectral Anomaly Detection via Structured Sparsity Plus Enhanced Low-Rankness
abstract
Hyperspectral anomaly detection (HAD), distinguishing anomalous pixels or subpixels from the background, has received increasing attention in recent years. Low-Rank Representation (LRR)-based methods have also been promoted rapidly for HAD, but they may encounter three challenges: (1) they adopted the nuclear norm as the convex approximation, yet a sub-optimal solution of the rank function; (2) they overlook the structured spatial correlation of anomalous pixels; (3) they fail to comprehensively explore the local structure details of the original background. To address these challenges, in this paper, we proposed the Structured Sparsity Plus Enhanced Low-Rank (S2ELR) method for HAD. Specifically, our S2ELR method adopts the weighted tensor Schatten-pnorm, acting as an enhanced approximation of the rank function than the tensor nuclear norm, and the structured sparse norm to characterize the low-rank properties of the background and the sparsity of the abnormal pixels, respectively. To preserve the local structural details, the position-based Laplace regularizer is accompanied. An iterative algorithm is derived from the popular alternating direction methods of multipliers. Compared to the existing state-of-the-art HAD methods, the experimental results have demonstrated the superiority of our proposed S2ELR method.
Yin-Ping Zhao, Yongyong Chen, Zhen Wang 0004, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Toward Complete-View and High-Level Pose-Based Gait Recognition
abstract
Model-based gait recognition methods usually adopt the pedestrian walking postures to identify human beings. However, existing methods did not explicitly resolve the large intra-class variance of human pose due to changes in camera view. In this paper, we propose a lower-upper generative adversarial network (LUGAN) to generate multi-view pose sequences for each single-view sample to reduce the cross-view variance. Based on the prior of camera imaging, we prove that the spatial coordinates between cross-view poses satisfy a linear transformation of a full-rank matrix. Hence, LUGAN employs the adversarial training to learn full-rank transformation matrices from the source pose and target views to obtain the target pose sequences. The generator of LUGAN is composed of graph convolutional (GCN) layers, fully connected (FC) layers and two-branch convolutional (CNN) layers: GCN layers and FC layers encode the source pose sequence and target view, then CNN layers take as input the encoded features to learn a lower triangular matrix and an upper one, finally the transformation matrix is formulated by multiplying the lower and upper triangular matrices. For the purpose of adversarial training, we develop a conditional discriminator that distinguishes whether the pose sequence is true or generated. Furthermore, to facilitate the high-level correlation learning, we propose a plug-and-play module, named multi-scale hypergraph convolution (HGC), to replace the spatial graph convolutional layer in baseline, which can simultaneously model the joint-level, part-level and body-level correlations. Extensive experiments on three large gait recognition datasets (i.e., CASIA-B, OUMVLP-Pose and NLPR) demonstrate that our method outperforms the baseline model by a large margin.
Honghu Pan, Yongyong Chen, Tingyang Xu, Yunqi He, Zhenyu He 0001
IEEE Trans. Inf. Forensics Secur.2
2023 Bi-Nuclear Tensor Schatten-p Norm Minimization for Multi-View Subspace Clustering
abstract
Multi-view subspace clustering aims to integrate the complementary information contained in different views to facilitate data representation. Currently, low-rank representation (LRR) serves as a benchmark method. However, we observe that these LRR-based methods would suffer from two issues: limited clustering performance and high computational cost since (1) they usually adopt the nuclear norm with biased estimation to explore the low-rank structures; (2) the singular value decomposition of large-scale matrices is inevitably involved. Moreover, LRR may not achieve low-rank properties in both intra-views and inter-views simultaneously. To address the above issues, this paper proposes the Bi-nuclear tensor Schatten- p norm minimization for multi-view subspace clustering (BTMSC). Specifically, BTMSC constructs a third-order tensor from the view dimension to explore the high-order correlation and the subspace structures of multi-view features. The Bi-Nuclear Quasi-Norm (BiN) factorization form of the Schatten- p norm is utilized to factorize the third-order tensor as the product of two small-scale third-order tensors, which not only captures the low-rank property of the third-order tensor but also improves the computational efficiency. Finally, an efficient alternating optimization algorithm is designed to solve the BTMSC model. Extensive experiments with ten datasets of texts and images illustrate the performance superiority of the proposed BTMSC method over state-of-the-art methods.
Shuqin Wang 0001, Zhiping Lin 0001, Qi Cao 0002, Yi-Gang Cen, Yongyong Chen
IEEE Trans. Image Process.5
2023 Multi-view Ensemble Clustering via Low-rank and Sparse Decomposition: From Matrix to Tensor
abstract
As a significant extension of classical clustering methods, ensemble clustering first generates multiple basic clusterings and then fuses them into one consensus partition by solving a problem concerning graph partition with respect to the co-association matrix. Although the collaborative cluster structure among basic clusterings can be well discovered by ensemble clustering, most advanced ensemble clustering utilizes the self-representation strategy with the constraint of low-rank to explore a shared consensus representation matrix in multiple views. However, they still encounter two challenges: (1) high computational cost caused by both the matrix inversion operation and singular value decomposition of large-scale square matrices; (2) less considerable attention on high-order correlation attributed to the pursue of the two-dimensional pair-wise relationship matrix. In this article, based on low-rank and sparse decomposition from both matrix and tensor perspectives, we propose two novel multi-view ensemble clustering methods, which tangibly decrease computational complexity. Specifically, our first method utilizes low-rank and sparse matrix decomposition to learn one common co-association matrix, while our last method constructs all co-association matrices into one third-order tensor to investigate the high-order correlation among multiple views by low-rank and sparse tensor decomposition. We adopt the alternating direction method of multipliers to solve two convex models by dividing them into several subproblems with closed-form solution. Experimental results on ten real-world datasets prove the effectiveness and efficiency of the proposed two multi-view ensemble clustering methods by comparing them with other advanced ensemble clustering methods.
Xuanqi Zhang, Qiangqiang Shen, Yongyong Chen, Zhongyun Hua, Jingyong Su
ACM Trans. Knowl. Discov. Data3
2023 Optimizing Spaced Repetition Schedule by Capturing the Dynamics of Memory
abstract
Spaced repetition, namely, learners review items in a given schedule, has been proven powerful for memorization and practice of skills. Most current spaced repetition methods focus on either predicting student recall or designing an optimal review schedule, thus omitting the integrity of the spaced repetition system. In this work, we propose a novel spaced repetition schedule framework by capturing the dynamics of memory, which alternates memory prediction and schedule optimization to improve the efficiency of learners’ reviews. First, the framework collects logs from students’ reviews and builds memory models with Markov property to capture the dynamics of memory. Then, the spaced repetition optimization is transformed a stochastic shortest path problem and solved via the value iteration method. We also construct a new benchmark dataset for spaced repetition, which is the first to contain time-series information during learners’ memorization. Experimental results on the collected data from the real world and the simulated environment demonstrate that the proposed approach reduces 64% error and 17% cost in predicting recall rates and optimizing schedules compared to several baselines. We have publicly released the dataset containing 220 million rows and codes used in this paper at:https://github.com/maimemo/SSP-MMC-Plus.
Jingyong Su, Junyao Ye, Liqiang Nie, Yilong Cao, Yongyong Chen
IEEE Trans. Knowl. Data Eng.5
2023 Unified Low-Rank Tensor Learning and Spectral Embedding for Multi-View Subspace Clustering
abstract
Multi-view subspace clustering aims to utilize the comprehensive information of multi-source features to aggregate data into multiple subspaces. Recently, low-rank tensor learning has been applied to multi-view subspace clustering, which explores high-order correlations of multi-view data and has achieved remarkable results. However, these existing methods have certain limitations: 1) The learning processes of low-rank tensor and label indicator matrix are independent. 2) Variable contributions of different views to the consistent clustering results are not discriminated. To handle these issues, we propose a unified framework that integrates low-rank tensor learning and spectral embedding (ULTLSE) for multi-view subspace clustering. Specifically, the proposed model adopts the tensor singular value decomposition (t-SVD) based tensor nuclear norm to encode the low-rank property of the self-representation tensor, and a label indicator matrix via spectral embedding is simultaneously exploited. To distinguish the importance of various views, we learn a quantifiable weighting coefficient for each view. An effective recursion optimization algorithm is also developed to address the proposed model. Finally, we conduct comprehensive experiments on eight real-world datasets with three categories. The experimental results indicate that the proposed ULTLSE is advanced over existing state-of-the-art clustering methods.
Lele Fu, Zhaoliang Chen, Yongyong Chen, Shiping Wang
IEEE Trans. Multim.3
2023 Bilateral Fast Low-Rank Representation With Equivalent Transformation for Subspace Clustering
abstract
In recent years, low-rank representation (LRR) has received increasing attention on subspace clustering. Due to inevitable matrix inversion and singular value decomposition in each iteration, however, most of existing LRR algorithms may suffer from high computational complexity, and hence can not cope with the large-scale sample data commendably. To overcome this problem, in this paper, we propose a bilateral fast low-rank representation (BFLRR), which has a linear time complexity with respect to the number of samples. Specifically, we introduce the equivalent transformation method to remove the null spaces of both the columns and rows of the coefficient matrix so that a hypercompact coefficient matrix can be learned. Furthermore, the proposed BFLRR is embedded into a distributed framework as DFC-BFLRR to make it more efficient, which utilizes a combination of the global and local projection matrices. Extensive experiments are carried out on real datasets, and the results testify that the proposed methods not only perform faster-computing speed but also obtain favorable clustering accuracy in comparison with the competing methods among large-scale sample data.
Qiangqiang Shen, Shuangyan Yi, Yongsheng Liang 0001, Yongyong Chen, Wei Liu 0065
IEEE Trans. Multim.4
2023 AMS-Net: Adaptive Multi-Scale Network for Image Compressive Sensing
abstract
Recently, deep convolutional neural networks have been applied to image compressive sensing (CS) to improve reconstruction quality while reducing computation cost. Existing deep learning-based CS methods can be divided into two classes: sampling image at single scale and sampling image across multiple scales. However, these existing methods treat the image low-frequency and high-frequency components equally, which is an obstruction to get a high reconstruction quality. This paper proposes an adaptive multi-scale image CS network in wavelet domain called AMS-Net, which fully exploits the different importance of image low-frequency and high-frequency components. First, the discrete wavelet transform is used to decompose an image into four sub-bands, namely the low-low (LL), low-high (LH), high-low (HL), and high-high (HH) sub-bands. Considering that the LL sub-band is more important to the final reconstruction quality, the AMS-Net allocates it a larger sampling ratio, while allocating the other three sub-bands a smaller one. Since different blocks in each sub-band have different sparsity, the sampling ratio is further allocated block-by-block within the four sub-bands. Then a dual-channel scalable sampling model is developed to adaptively sample the LL and the other three sub-bands at arbitrary sampling ratios. Finally, by unfolding the iterative reconstruction process of the traditional multi-scale block CS algorithm, we construct a multi-stage reconstruction model to utilize multi-scale features for further improving the reconstruction quality. Experimental results demonstrate that the proposed model outperforms both the traditional and state-of-the-art deep learning-based methods.
Zhongyun Hua, Yuanman Li, Yongyong Chen, Yicong Zhou
IEEE Trans. Multim.4
2023 Structured anchor-inferred graph learning for universal incomplete multi-view clustering
Wenjue He, Zheng Zhang 0006, Yongyong Chen, Jie Wen 0001
World Wide Web (WWW)3
2022 Correntropy-Induced Tensor Learning for Multi-view Subspace Clustering
abstract
Using some specific optimization problems with specific regularizers, multi-view subspace clustering has achieved better performance over single-view subspace clustering. However, they simply assume the noise obeys the Gaussian distribution only, and thus the dataset with non-Gaussian noise or outliers may not be accurately clustered. To address this issue, this paper proposes a novel correntropy-induced tensor learning method for multi-view subspace clustering (CTMSC). Specifically, CTMSC adopts the correntropy-induced metric to substitute the traditional mean square error (MSE) to handle non-Gaussian noise or outliers. Furthermore, the proposed objective function is optimized using an alternating direction method of multipliers with the aid of half-quadratic technology in the form of multiplication. Extensive experimental results on various real-world datasets demonstrate the effectiveness of the proposed method by comparing several state-of-the-art multi-view subspace clustering methods.
Yongyong Chen, Shuqin Wang 0001, Jingyong Su, Junxin Chen 0001
ICDM1
2022 Deep Contrastive Multi-view Subspace Clustering
Yongyong Chen, Zhongyun Hua
ICONIP (4)2
2022 Nonconvex low-rank and sparse tensor representation for multi-view subspace clustering
Shuqin Wang 0001, Yongyong Chen, Yi-Gang Cen, Linna Zhang, Hengyou Wang, Viacheslav V. Voronin
Appl. Intell.2
2022 Frobenius norm-regularized robust graph learning for multi-view subspace clustering
Shuqin Wang 0001, Yongyong Chen, Guoqing Chao
Appl. Intell.2
2022 TPpred-ATMV: therapeutic peptide prediction by adaptive multi-view tensor learning model
abstract
MOTIVATION: Therapeutic peptide prediction is important for the discovery of efficient therapeutic peptides and drug development. Researchers have developed several computational methods to identify different therapeutic peptide types. However, these computational methods focus on identifying some specific types of therapeutic peptides, failing to predict the comprehensive types of therapeutic peptides. Moreover, it is still challenging to utilize different properties to predict the therapeutic peptides. RESULTS: In this study, an adaptive multi-view based on the tensor learning framework TPpred-ATMV is proposed for predicting different types of therapeutic peptides. TPpred-ATMV constructs the class and probability information based on various sequence features. We constructed the latent subspace among the multi-view features and constructed an auto-weighted multi-view tensor learning model to utilize the high correlation based on the multi-view features. Experimental results showed that the TPpred-ATMV is better than or highly comparable with the other state-of-the-art methods for predicting eight types of therapeutic peptides. AVAILABILITY AND IMPLEMENTATION: The code of TPpred-ATMV is accessed at: https://github.com/cokeyk/TPpred-ATMV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ke Yan 0003, Hongwu Lv, Yongyong Chen, Hao Wu 0066, Bin Liu 0014
Bioinform.4
2022 Log-based sparse nonnegative matrix factorization for data representation
Chong Peng 0001, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
Knowl. Based Syst.3
2022 Preserving bilateral view structural information for subspace clustering
Chong Peng 0001, Yongyong Chen, Chenglizhao Chen, Zhao Kang 0001, Li Guo 0016, Qiang Shawn Cheng
Knowl. Based Syst.3
2022 Weighted Schatten p-norm minimization with logarithmic constraint for subspace clustering
Qiangqiang Shen, Yongyong Chen, Yongsheng Liang 0001, Shuangyan Yi, Wei Liu 0065
Signal Process.2
2022 Low-Rank Tensor Graph Learning for Multi-View Subspace Clustering
abstract
Graph and subspace clustering methods have become the mainstream of multi-view clustering due to their promising performance. However, (1) since graph clustering methods learn graphs directly from the raw data, when the raw data is distorted by noise and outliers, their performance may seriously decrease; (2) subspace clustering methods use a “two-step” strategy to learn the representation and affinity matrix independently, and thus may fail to explore their high correlation. To address these issues, we propose a novel multi-view clustering method via learning aLow-RankTensorGraph (LRTG). Different from subspace clustering methods, LRTG simultaneously learns the representation and affinity matrix in a single step to preserve their correlation. We apply Tucker decomposition and$l_{2,1}$-norm to the LRTG model to alleviate noise and outliers for learning a “clean” representation. LRTG then learns the affinity matrix from this “clean” representation. Additionally, an adaptive neighbor scheme is proposed to find the$K$largest entries of the affinity matrix to form a flexible graph for clustering. An effective optimization algorithm is designed to solve the LRTG model based on the alternating direction method of multipliers. Extensive experiments on different clustering tasks demonstrate the effectiveness and superiority of LRTG over seventeen state-of-the-art clustering methods.
Yongyong Chen, Xiaolin Xiao, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2022 TCDesc: Learning Topology Consistent Descriptors for Image Matching
abstract
The triplet loss is widely used in learning the local descriptors for image matching. However, existing triplet loss-based methods, like HardNet and DSM, employ the point-to-point distance metric, which neglects the neighborhood information of descriptors. Considering the fact that local neighborhood structures of matching descriptors would be similar under the ideal condition, this paper aims to learn the neighborhood topology-consistent descriptors (TCDesc). To this end, we first propose the linear combination weight as the topology weight to depict the neighborhood topology for each descriptor, where the difference between the center descriptor and the linear combination of its neighbors is minimized. For the global comparison, we then define a global topology vector by using the local topology weights. Next, beyond the Euclidean distance, we define a topology distance with the topology vectors to indicate the topological difference between the matching descriptors. Furthermore, we propose an adaptive weighting strategy to jointly minimize the topology distance and Euclidean distance in triplet loss. Experimental results on four widely-used datasets, i.e., UBC PhotoTourism, HPatches, W1BS and Oxford, demonstrate that our method can effectively improve the performance of both HardNet and DSM.
Honghu Pan, Yongyong Chen, Zhenyu He 0001, Fanyang Meng, Nana Fan
IEEE Trans. Circuits Syst. Video Technol.2
2022 Hyperspectral Image Denoising Using Nonconvex Local Low-Rank and Sparse Separation With Spatial-Spectral Total Variation Regularization
abstract
In this paper, we propose a novel nonconvex approach to robust principal component analysis for HSI denoising, which focuses on simultaneously developing more accurate approximations to both rank and column-wise sparsity for the low-rank and sparse components, respectively. In particular, the new method adopts the log-determinant rank approximation and a novell2,lognorm, to restrict the local low-rank or column-wisely sparse properties for the component matrices, respectively. For thel2,log-regularized shrinkage problem, we develop an efficient, closed-form solution, which is namedl2,log-shrinkage operator. The new regularization and the corresponding operator can be generally used in other problems that require column-wise sparsity. Moreover, we impose the spatial-spectral total variation regularization in the log-based nonconvex RPCA model, which enhances the global piece-wise smoothness and spectral consistency from the spatial and spectral views in the recovered HSI. Extensive experiments on both simulated and real HSIs demonstrate the effectiveness of the proposed method in denoising HSIs.
Chong Peng 0001, Kehan Kang, Yongyong Chen, Xinxing Wu, Andrew Cheng, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
IEEE Trans. Geosci. Remote. Sens.4
2022 DDCNN: A Deep Learning Model for AF Detection From a Single-Lead Short ECG Signal
abstract
With the popularity of the wireless body sensor network, real-time and continuous collection of single-lead electrocardiogram (ECG) data becomes possible in a convenient way. Data mining from the collected single-lead ECG waves has therefore aroused extensive attention worldwide, where early detection of atrial fibrillation (AF) is a hot research topic. In this paper, a two-channel convolutional neural network combined with a data augmentation method is proposed to detect AF from single-lead short ECG recordings. It consists of three modules, the first module denoises the raw ECG signals and produces 9-s ECG signals and heart rate (HR) values. Then, the ECG signals and HR rate values are fed into the convolutional layers for feature extraction, followed by three fully connected layers to perform the classification. The data augmentation method is used to generate synthetic signals to enlarge the training set and increase the diversity of the single-lead ECG signals. Validation experiments and the comparison with state-of-the-art studies demonstrate the effectiveness and advantages of the proposed method.
Zhaocheng Yu, Junxin Chen 0001, Yu Liu 0035, Yongyong Chen, Tingting Wang 0006, Robert M. Nowak, Zhihan Lyu
IEEE J. Biomed. Health Informatics4
2022 Self-Paced Enhanced Low-Rank Tensor Kernelized Multi-View Subspace Clustering
abstract
This paper addresses the multi-view subspace clustering problem and proposes the self-paced enhanced low-rank tensor kernelized multi-view subspace clustering (SETKMC) method, which is based on two motivations: (1) singular values of the representations and multiple instances should be treated differently. The reasons are that larger singular values of the representations usually quantify the major information and should be less penalized; samples with different degrees of noise may have various reliability for clustering. (2) many existing methods may cause the degraded performance when multi-view features reside in different nonlinear subspaces. This is because they usually assumed that multiple features lie within the union of several linear subspaces. SETKMC integrates the nonconvex tensor norm, self-paced learning, and kernel trick into a unified model for multi-view subspace clustering. The nonconvex tensor norm imposes different weights on different singular values. The self-paced learning gradually involves instances from more reliable to less reliable ones while the kernel trick aims to handle the multi-view data in nonlinear subspaces. One iterative algorithm is proposed based on the alternating direction method of multipliers. Extensive results on seven real-world datasets show the effectiveness of the proposed SETKMC compared to fifteen state-of-the-art multi-view clustering methods.
Yongyong Chen, Shuqin Wang 0001, Xiaolin Xiao, Youfa Liu, Zhongyun Hua, Yicong Zhou
IEEE Trans. Multim.1
2022 Adaptive Transition Probability Matrix Learning for Multiview Spectral Clustering
abstract
Multiview clustering as an important unsupervised method has been gathering a great deal of attention. However, most multiview clustering methods exploit theself-representation propertyto capture the relationship among data, resulting in high computation cost in calculating the self-representation coefficients. In addition, they usually employ different regularizers to learn the representation tensor or matrix from which a transition probability matrix is constructed in a separate step, such as the one proposed by Wuet al.. Thus, an optimal transition probability matrix cannot be guaranteed. To solve these issues, we propose a unified model for multiview spectral clustering by directly learning an adaptive transition probability matrix (MCA2M), rather than an individual representation matrix of each view. Different from the one proposed by Wuet al., MCA2M utilizes the one-step strategy to directly learn the transition probability matrix under the robust principal component analysis framework. Unlike existing methods using the absolute symmetrization operation to guarantee the nonnegativity and symmetry of the affinity matrix, the transition probability matrix learned from MCA2M is nonnegative and symmetric without any postprocessing. An alternating optimization algorithm is designed based on the efficient alternating direction method of multipliers. Extensive experiments on several real-world databases demonstrate that the proposed method outperforms the state-of-the-art methods.
Yongyong Chen, Xiaolin Xiao, Zhongyun Hua, Yicong Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2022 Two-Dimensional Parametric Polynomial Chaotic System
abstract
When used in engineering applications, most existing chaotic systems may have many disadvantages, including discontinuous chaotic parameter ranges, lack of robust chaos, and easy occurrence of chaos degradation. In this article, we propose a two-dimensional (2-D) parametric polynomial chaotic system (2D-PPCS) as a general system that can yield many 2-D chaotic maps with different exponent coefficient settings. The 2D-PPCS initializes two parametric polynomials and then applies modular chaotification to the polynomials. Setting different control parameters allows the 2D-PPCS to customize its Lyapunov exponents in order to obtain robust chaos and behaviors with desired complexity. Our theoretical analysis demonstrates the robust chaotic behavior of the 2D-PPCS. Two illustrative examples are provided and tested based on numeral experiments to verify the effectiveness of the 2D-PPCS. A chaos-based pseudorandom number generator is also developed to illustrate the applications of the 2D-PPCS. The experimental results demonstrate that these examples of the 2D-PPCS can achieve robust and desired chaos, have better performance, and generate higher randomness pseudorandom numbers than some representative 2-D chaotic maps.
Zhongyun Hua, Yongyong Chen, Han Bao 0001, Yicong Zhou
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Hyperspectral Image Denoising With Log-Based Robust PCA
abstract
It is a challenging task to remove heavy and mixed types of noise from Hyperspectral images (HSIs). In this paper, we propose a novel nonconvex approach to RPCA for HSI denoising, which adopts the log-determinant rank approximation and a novel $\ell_{2,\text{l}\text{o}\text{g}}$ norm, to restrict the low-rank or column-wise sparse properties for the component matrices, respectively. For the $\ell_{2,\text{l}\text{o}\text{g}}$-regularized shrinkage problem, we develop an efficient, closed-form solution, which is named $\ell_{2,\text{l}\text{o}\text{g}}$-shrinkage operator, which can be generally used in other problems. Extensive experiments on both simulated and real HSIs demonstrate the effectiveness of the proposed method in denoising HSIs.
Yongyong Chen, Qiang Shawn Cheng, Chong Peng 0001
ICIP3
2021 Low-Rank And Sparse Tensor Representation For Multi-View Subspace Clustering
abstract
Learning an effective affinity matrix as the input of spectral clustering to achieve promising multi-view clustering is a key issue of subspace clustering. In this paper, we propose a low-rank and sparse tensor representation (LRSTR) method that learns the affinity matrix through a self-representation tensor and retains the similarity information of the view dimensions for multi-view subspace clustering. Specifically, the proposed LRSTR method imposes the tensor nuclear norm and tensor sparse constraints on self-representation tensor to characterize the relationship between views. The optimization model is solved under the framework of alternating direction method of multiplier. Experimental results on four datasets show that the proposed LRSTR method is better than several state-of-the-art methods.
Shuqin Wang 0001, Yongyong Chen, Yigang Ce, Linna Zhang, Viacheslav V. Voronin
ICIP2
2021 Partial Tubal Nuclear Norm Regularized Multi-view Learning
abstract
Multi-view clustering and multi-view dimension reduction explore ubiquitous and complementary information between multiple features to enhance the clustering, recognition performance. However, multi-view clustering and multi-view dimension reduction are treated independently, ignoring the underlying correlations between them. In addition, previous methods mainly focus on using the tensor nuclear norm for low-rank representation to explore the high correlation of multi-view features, which often causes the estimation bias of the tensor rank. To overcome these limitations, we propose the partial tubal nuclear norm regularized multi-view learning (PTN2ML) method, in which the partial tubal nuclear norm as a non-convex surrogate of the tensor tubal multi-rank, only minimizes the partial sum of the smaller tubal singular values to preserve the low-rank property of the self-representation tensor. PTN2ML pursues the latent representation from the projection space rather than from the input space to reveal the structural consensus and suppress the disturbance of noisy data. The proposed method can be efficiently optimized by the alternating direction method of multipliers. Extensive experiments, including multi-view clustering and multi-view dimension reduction substantiate the superiority of the proposed methods beyond state-of-the-arts.
Yongyong Chen, Shuqin Wang 0001, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou
ACM Multimedia1
2021 Learning discriminative representation for image classification
Chong Peng 0001, Zhao Kang 0001, Yongyong Chen, Chenglizhao Chen, Qiang Shawn Cheng
Knowl. Based Syst.5
2021 Error-robust low-rank tensor approximation for multi-view clustering
Shuqin Wang 0001, Yongyong Chen, Yi Jin 0001, Yi-Gang Cen, Yidong Li, Linna Zhang
Knowl. Based Syst.2
2021 Generalized Nonconvex Low-Rank Tensor Approximation for Multi-View Subspace Clustering
abstract
The low-rank tensor representation (LRTR) has become an emerging research direction to boost the multi-view clustering performance. This is because LRTR utilizes not only the pairwise relation between data points, but also the view relation of multiple views. However, there is one significant challenge: LRTR uses the tensor nuclear norm as the convex approximation but provides a biased estimation of the tensor rank function. To address this limitation, we propose the generalized nonconvex low-rank tensor approximation (GNLTA) for multi-view subspace clustering. Instead of the pairwise correlation, GNLTA adopts the low-rank tensor approximation to capture the high-order correlation among multiple views and proposes the generalized nonconvex low-rank tensor norm to well consider the physical meanings of different singular values. We develop a unified solver to solve the GNLTA model and prove that under mild conditions, any accumulation point is a stationary point of GNLTA. Extensive experiments on seven commonly used benchmark databases have demonstrated that the proposed GNLTA achieves better clustering performance over state-of-the-art methods.
Yongyong Chen, Shuqin Wang 0001, Chong Peng 0001, Zhongyun Hua, Yicong Zhou
IEEE Trans. Image Process.1
2021 Low-Rank Preserving t-Linear Projection for Robust Image Feature Extraction
abstract
As the cornerstone for joint dimension reduction and feature extraction, extensive linear projection algorithms were proposed to fit various requirements. When being applied to image data, however, existing methods suffer from representation deficiency since the multi-way structure of the data is (partially) neglected. To solve this problem, we propose a novel Low-Rank Preserving t-Linear Projection (LRP-tP) model that preserves the intrinsic structure of the image data using t-product-based operations. The proposed model advances in four aspects: 1) LRP-tP learns the t-linear projection directly from the tensorial dataset so as to exploit the correlation among the multi-way data structure simultaneously; 2) to cope with the widely spread data errors, e.g., noise and corruptions, the robustness of LRP-tP is enhanced via self-representation learning; 3) LRP-tP is endowed with good discriminative ability by integrating the empirical classification error into the learning procedure; 4) an adaptive graph considering the similarity and locality of the data is jointly learned to precisely portray the data affinity. We devise an efficient algorithm to solve the proposed LRP-tP model using the alternating direction method of multipliers. Extensive experiments on image feature extraction have demonstrated the superiority of LRP-tP compared to the state-of-the-arts.
Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou
IEEE Trans. Image Process.2
2021 Prior Knowledge Regularized Multiview Self-Representation and its Applications
abstract
To learn the self-representation matrices/tensor that encodes the intrinsic structure of the data, existing multiview self-representation models consider only the multiview features and, thus, impose equal membership preference across samples. However, this is inappropriate in real scenarios since the prior knowledge, e.g., explicit labels, semantic similarities, and weak-domain cues, can provide useful insights into the underlying relationship of samples. Based on this observation, this article proposes a prior knowledge regularized multiview self-representation (P-MVSR) model, in which the prior knowledge, multiview features, and high-order cross-view correlation are jointly considered to obtain an accurate self-representation tensor. The general concept of "prior knowledge" is defined as the complement of multiview features, and the core of P-MVSR is to take advantage of the membership preference, which is derived from the prior knowledge, to purify and refine the discovered membership of the data. Moreover, P-MVSR adopts the same optimization procedure to handle different prior knowledge and, thus, provides a unified framework for weakly supervised clustering and semisupervised classification. Extensive experiments on real-world databases demonstrate the effectiveness of the proposed P-MVSR model.
Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2020 Robust principal component analysis: A factorization-based approach with linear complexity
Chong Peng 0001, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng
Inf. Sci.2
2020 Graph-regularized least squares regression for multi-view subspace clustering
Yongyong Chen, Shuqin Wang 0001, Fangying Zheng, Yi-Gang Cen
Knowl. Based Syst.1
2020 Multi-view subspace clustering via simultaneously learning the representation tensor and affinity matrix
Yongyong Chen, Xiaolin Xiao, Yicong Zhou
Pattern Recognit.1
2020 Noninteractive Lightweight Privacy-Preserving Auditing on Images in Mobile Crowdsourcing Networks
abstract
To determine whether images on the crowdsourcing server meet the mobile user’s requirement, an auditing protocol is desired to check these images. However, before paying for images, the mobile user typically cannot download them for checking. Moreover, since mobiles are usually low-power devices and the crowdsourcing server has to handle a large number of mobile users, the auditing protocol should be lightweight. To address the above security and efficiency issues, we propose a novel noninteractive lightweight privacy-preserving auditing protocol on images in mobile crowdsourcing networks, called NLPAS. Since NLPAS allows the mobile user to check images on the crowdsourcing server without downloading them, the newly designed protocol can provide privacy protection for these images. At the same time, NLPAS uses the binary convolutional neural network for extracting features from images and designs a novel privacy-preserving Hamming distance computation algorithm for determining whether these images on the crowdsourcing server meet the mobile user’s requirement. Since these two techniques are both lightweight, NLPAS can audit images on the crowdsourcing server in a privacy-preserving manner while still enjoying high efficiency. Experimental results show that NLPAS is feasible for real-world applications.
Changsheng Wan, Yongyong Chen
Secur. Commun. Networks5
2020 Hyperspectral image denoising by total variation-regularized bilinear factorization
Yongyong Chen, Jiaxue Li, Yicong Zhou
Signal Process.1
2020 Low-Rank Quaternion Approximation for Color Image Processing
abstract
Low-rank matrix approximation (LRMA)-based methods have made a great success for grayscale image processing. When handling color images, LRMA either restores each color channel independently using the monochromatic model or processes the concatenation of three color channels using the concatenation model. However, these two schemes may not make full use of the high correlation among RGB channels. To address this issue, we propose a novel low-rank quaternion approximation (LRQA) model. It contains two major components: first, instead of modeling a color image pixel as a scalar in conventional sparse representation and LRMA-based methods, the color image is encoded as a pure quaternion matrix, such that the cross-channel correlation of color channels can be well exploited; second, LRQA imposes the low-rank constraint on the constructed quaternion matrix. To better estimate the singular values of the underlying low-rank quaternion matrix from its noisy observation, a general model for LRQA is proposed based on several nonconvex functions. Extensive evaluations for color image denoising and inpainting tasks verify that LRQA achieves better performance over several state-of-the-art sparse representation and LRMA-based methods in terms of both quantitative metrics and visual quality.
Yongyong Chen, Xiaolin Xiao, Yicong Zhou
IEEE Trans. Image Process.1
2020 2D Quaternion Sparse Discriminant Analysis
abstract
Linear discriminant analysis has been incorporated with various representations and measurements for dimension reduction and feature extraction. In this paper, we propose two-dimensional quaternion sparse discriminant analysis (2D-QSDA) that meets the requirements of representing RGB and RGB-D images. 2D-QSDA advances in three aspects: 1) including sparse regularization, 2D-QSDA relies only on the important variables, and thus shows good generalization ability to the out-of-sample data which are unseen during the training phase; 2) benefited from quaternion representation, 2D-QSDA well preserves the high order correlation among different image channels and provides a unified approach to extract features from RGB and RGB-D images; 3) the spatial structure of the input images is retained via the matrix-based processing. We tackle the constrained trace ratio problem of 2D-QSDA by solving a corresponding constrained trace difference problem, which is then transformed into a quaternion sparse regression (QSR) model. Afterward, we reformulate the QSR model to an equivalent complex form to avoid the processing of the complicated structure of quaternions. A nested iterative algorithm is designed to learn the solution of 2D-QSDA in the complex space and then we convert this solution back to the quaternion domain. To improve the separability of 2D-QSDA, we further propose 2D-QSDAw using the weighted pairwise between-class distances. Extensive experiments on RGB and RGB-D databases demonstrate the effectiveness of 2D-QSDA and 2D-QSDAw compared with peer competitors.
Xiaolin Xiao, Yongyong Chen, Yue-Jiao Gong, Yicong Zhou
IEEE Trans. Image Process.2
2020 Jointly Learning Kernel Representation Tensor and Affinity Matrix for Multi-View Clustering
abstract
Multi-view clustering refers to the task of partitioning numerous unlabeled multimedia data into several distinct clusters using multiple features. In this paper, we propose a novel nonlinear method called joint learning multi-view clustering (JLMVC) to jointly learn kernel representation tensor and affinity matrix. The proposed JLMVC has three advantages: (1) unlike existing low-rank representation-based multi-view clustering methods that learn the representation tensor and affinity matrix in two separate steps, JLMVC jointly learns them both; (2) using the “kernel trick,” JLMVC can handle nonlinear data structures for various real applications; and (3) different from most existing methods that treat representations of all views equally, JLMVC automatically learns a reasonable weight for each view. Based on the alternating direction method of multipliers, an effective algorithm is designed to solve the proposed model. Extensive experiments on eight multimedia datasets demonstrate the superiority of the proposed JLMVC over state-of-the-art methods.
Yongyong Chen, Xiaolin Xiao, Yicong Zhou
IEEE Trans. Multim.1
2019 Multi-view Clustering via Simultaneously Learning Graph Regularized Low-Rank Tensor Representation and Affinity Matrix
abstract
Low-rank tensor representation-based multi-view clustering has become an efficient method for data clustering due to the robustness to noise and the preservation of the high order correlation. However, existing algorithms may suffer from two common problems: (1) the local view-specific geometrical structures and the various importance of features in different views are neglected; (2) the low-rank representation tensor and the affinity matrix are learned separately. To address these issues, we propose a novel framework to learn the Graph regularized Low-rank Tensor representation and the Affinity matrix (GLTA) in a unified manner. Besides, the manifold regularization is exploited to preserve the view-specific geometrical structures, and the various importance of different features is automatically calculated when constructing the final affinity matrix. An efficient algorithm is designed to solve GLTA using the augmented Lagrangian multiplier. Extensive experiments on six real datasets demonstrate the superiority of GLTA over the state-of-the-arts.
Yongyong Chen, Xiaolin Xiao, Yicong Zhou
ICME1
2018 Robust Principal Component Analysis with Matrix Factorization
abstract
Traditional robust principle component analysis (RPCA) has a high computational cost because RPCA needs to calculate the singular value decomposition of large matrices. To address this issue, this paper proposes a matrix-factorization-based RPCA (MFRPCA) model. MFRPCA has high computation efficiency while improving the robustness and flexibility of traditional RPCA using a non-convex low-rank approximation. Experiment results on challenging datasets demonstrate superior performance of MFRPCA compared with several advanced low-rank reconstruction methods.
Yongyong Chen, Yicong Zhou
ICASSP1
2018 Total Variation Regularized Low-Rank Tensor Approximation for Color Image Denoising
abstract
Existing approaches for low-rank approximation either need a rank prior or ignore the spatial smooth characteristic of a color image. To overcome these drawbacks, we propose a total variation regularized low-rank tensor approximation model for color image denoising. The model integrates the strong low-rank prior into a tensor-SVD framework, and introduces the hyper total variation to model the spatial smooth structure of images. Using the alternating direction method of multipliers, we propose a simple algorithm to solve our model. Extensive results on simulated and real noisy color images demonstrate the better performance of the proposed method against state-of-the-art denoising methods.
Yongyong Chen, Yicong Zhou
SMC1
2017 Denoising of Hyperspectral Image Using Low-Rank Matrix Factorization
abstract
Restoration of hyperspectral images (HSIs) is a challenging task, owing to the reason that images are inevitably contaminated by a mixture of noise, including Gaussian noise, impulse noise, dead lines, and stripes, during their acquisition process. Recently, HSI denoising approaches based on low-rank matrix approximation have become an active research field in remote sensing and have achieved state-of-the-art performance. These approaches, however, unavoidably require to calculate full or partial singular value decomposition of large matrices, leading to the relatively high computational cost and limiting their flexibility. To address this issue, this letter proposes a method exploiting a low-rank matrix factorization scheme, in which the associated robust principal component analysis is solved by the matrix factorization of the low-rank component. Our method needs only an upper bound of the rank of the underlying low-rank matrix rather than the precise value. The experimental results on the simulated and real data sets demonstrate the performance of our method by removing the mixed noise and recovering the severely contaminated images.
Fei Xu 0005, Yongyong Chen, Chong Peng 0001, Yongli Wang 0004, Guoping He
IEEE Geosci. Remote. Sens. Lett.2
2017 Image Projection Ridge Regression for Subspace Clustering
abstract
Subspace clustering methods have been widely studied recently. When the inputs are two-dimensional (2-D) data, existing subspace clustering methods usually convert them into vectors, which severely damages inherent structures and relationships from original data. In this letter, we propose a novel subspace clustering method for 2-D data. It directly uses 2-D data as inputs such that the learning of representations benefits from inherent structures and relationships of the data. It simultaneously seeks image projection and representation coefficients such that they mutually enhance each other and lead to powerful data representations. An efficient algorithm is developed to solve the proposed objective function with provable decreasing and convergence property. Extensive experimental results verify the effectiveness of the new method.
Chong Peng 0001, Zhao Kang 0001, Fei Xu 0005, Yongyong Chen, Qiang Shawn Cheng
IEEE Signal Process. Lett.4
2017 Denoising of Hyperspectral Images Using Nonconvex Low Rank Matrix Approximation
abstract
Hyperspectral image (HSI) denoising is challenging not only because of the difficulty in preserving both spectral and spatial structures simultaneously, but also due to the requirement of removing various noises, which are often mixed together. In this paper, we present a nonconvex low rank matrix approximation (NonLRMA) model and the corresponding HSI denoising method by reformulating the approximation problem using nonconvex regularizer instead of the traditional nuclear norm, resulting in a tighter approximation of the original sparsity-regularised rank function. NonLRMA aims to decompose the degraded HSI, represented in the form of a matrix, into a low rank component and a sparse term with a more robust and less biased formulation. In addition, we develop an iterative algorithm based on the augmented Lagrangian multipliers method and derive the closed-form solution of the resulting subproblems benefiting from the special property of the nonconvex surrogate function. We prove that our iterative optimization converges easily. Extensive experiments on both simulated and real HSIs indicate that our approach can not only suppress noise in both severely and slightly noised bands but also preserve large-scale image structures and small-scale details well. Comparisons against state-of-the-art LRMA-based HSI denoising approaches show our superior performance.
Yongyong Chen, Yanwen Guo 0001, Yongli Wang 0004, Chong Peng 0001, Guoping He
IEEE Trans. Geosci. Remote. Sens.1