EDBT 2026 Demo / reviewers in the wild / expert
Xiaorong Pu
dblp:74/6232
· DBLP profile ↗
60ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0001-7387-7194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 26 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Views Attention Fusion of Granular-ball Fuzzy Representations Split for Improved Multi-view ClusteringabstractMulti-View Clustering (MVC) is a pivotal multi-view learning paradigm widely adopted across various fields. Despite recent advances, existing methods primarily focus on enhancing the performance of fused multi-view representation, often neglecting the issue of Representation Degradation (RD) arising from discrepancies in the intrinsic quality of different views. To address the limitations, we propose a novel Granular-ball Fuzzy Split and Attention Fusion (GFSAF) learning, which leverages the nature of granular-ball to extract mutual and complementary representation separately. Meanwhile, the proposed method introduces an attention variant for fused representations to mitigate the RD issue. GFSAF mainly consists of two training stages: Split-Extract Stage and Views-Fusion Stage. Specifically, we design a novel Granular-ball Fuzzy Contrastive Learning to extract mutual representation, and introduce Noise Stripping Loss to reduce the influence of noise for complementary representation. Then, a novel multi-head Cross Views Attention is proposed to employ attention mechanism from multi-view perspectives for comprehensive fused representations. Experimental results on eight databases demonstrate that our GFSAF achieves superior performance compared to several state-of-the-art MVC methods. Shuaiyu Liu, Jie Xu 0044, Yazhou Ren 0001, Yang Yang 0002, Xiaorong Pu, Guoyin Wang 0001 |
AAAI | 6 |
| 2026 | Topology-Aware Vision Transformers for Enhanced Scene RecognitionabstractScene recognition (SR) is a fundamental task in computer vision (CV). In recent years, Transformer-based methods have achieved remarkable success in scene recognition tasks. Most existing approaches primarily rely on visual features, while failing to effectively model the structural relationships within scenes, which are crucial for accurate scene recognition. To this end, we propose Topology Attention Network for Scene Recognition (TANSR), an innovative method that leverages topological relationships from graphs to guide scene recognition. Specifically, Graph Attention Mask Generation Network (GAMGN) generates topology-aware masks from graph representations constructed by Graph Generation Module (GGM) and integrates them with patch embeddings by Topology Attention Guidance (TAG), enabling the transformer's attention mechanism to incorporate topological information. Furthermore, we introduce an innovative attention-driven multimodal fusion strategy that integrates graph-derived topological cues with visual patch embeddings, substantially enhancing the transformer’s capability to capture topological information and improving performance in complex scene recognition tasks. We evaluate TANSR on the benchmarks MIT-67, Scene-15 and SUN397, where it achieves consistent state-of-the-art (SOTA) performance, including 98.58% accuracy on MIT-67. Yunxi Wang, Shuaiyu Liu, Qiling Li, Yazhou Ren 0001, Xiaorong Pu |
AAAI | 5 |
| 2025 | VSNet: Focusing on the Linguistic Characteristics of Sign LanguageabstractSign language is a visual language expressed through complex movements of the upper body. The human skeleton plays a critical role in sign language recognition due to its good separation from the video background. However, mainstream skeleton-based sign language recognition models often overly focus on the natural connections between joints, treating sign language as ordinary human movements, which neglects its linguistic characteristics. We believe that just as letters form words, each sign language gloss can also be decomposed into smaller visual symbols. To fully harness the potential of skeleton data, this paper proposes a novel joint fusion strategy and a visual symbol attention model. Specifically, we first input the complete set of skeletal joints, and after dynamically exchanging joint information, we discard the parts with the weakest connections to other joints, resulting in a fused, simplified skeleton. Then, we group the joints most likely to express the same visual symbol and discuss the joint movements within each group separately. To validate the superiority of our method, we conduct extensive experiments on multiple public benchmark datasets. The results show that, without complex pre-training, we still achieve new state-of-the-art performance. The code is available at https://github.com/atinyboy/VSNet. Xinyue Chen 0004, Xiaorong Pu, Yazhou Ren 0001 |
CVPR | 4 |
| 2025 | Multi-View Graph Clustering via Node-Guided Contrastive EncodingabstractMulti-view clustering has gained significant attention for integrating multi-view information in multimedia applications. With the growing complexity of graph data, multi-view graph clustering (MVGC) has become increasingly important. Existing methods primarily use Graph Neural Networks (GNNs) to encode structural and feature information, but applying GNNs within contrastive learning poses specific challenges, such as integrating graph data with node features and handling both homophilic and heterophilic graphs. To address these challenges, this paper introduces Node-Guided Contrastive Encoding (NGCE), a novel MVGC approach that leverages node features to guide embedding generation. NGCE enhances compatibility with GNN filtering, effectively integrates homophilic and heterophilic information, and strengthens contrastive learning across views. Extensive experiments demonstrate its robust performance on six homophilic and heterophilic multi-view benchmark datasets. Yazhou Ren 0001, Junlong Ke, Zichen Wen, Yang Yang 0002, Xiaorong Pu, Lifang He 0001 |
ICML | 6 |
| 2025 | An Effective and Secure Federated Multi-View Clustering Method with Information-Theoretic PerspectiveabstractRecently, federated multi-view clustering (FedMVC) has gained attention for its ability to mine complementary clustering structures from multiple clients without exposing private data. Existing methods mainly focus on addressing the feature heterogeneity problem brought by views on different clients and mitigating it using shared client information. Although these methods have achieved performance improvements, the information they choose to share, such as model parameters or intermediate outputs, inevitably raises privacy concerns. In this paper, we propose an Effective and Secure Federated Multi-view Clustering method, ESFMC, to alleviate the dilemma between privacy protection and performance improvement. This method leverages the information-theoretic perspective to split the features extracted locally by clients, retaining sensitive information locally and only sharing features that are highly relevant to the task. This can be viewed as a form of privacy-preserving information sharing, reducing privacy risks for clients while ensuring that the server can mine high-quality global clustering structures. Theoretical analysis and extensive experiments demonstrate that the proposed method more effectively mitigates the trade-off between privacy protection and performance improvement compared to state-of-the-art methods. Xinyue Chen 0004, Jinfeng Peng, Xiaorong Pu, Yang Yang 0002, Yazhou Ren 0001 |
ICML | 4 |
| 2025 | Lightweight Medical Image Restoration via Integrating Reliable Lesion-Semantic Driven PriorabstractMedical image restoration tasks aim to recover high-quality images from degraded observations, exhibiting emergent desires in many clinical scenarios, such as low-dose CT image denoising, MRI super-resolution, and MRI artifact removal. Despite the success achieved by existing deep learning-based restoration methods with sophisticated modules, they struggle with rendering computationally-efficient reconstruction results. Moreover, they usually ignore the reliability of the restoration results, which is much more urgent in medical systems. To alleviate these issues, we present LRformer, a Lightweight Transformer-based method via Reliability-guided learning in the frequency domain. Specifically, inspired by the uncertainty quantification in Bayesian neural networks (BNNs), we develop a Reliable Lesion-Semantic Prior Producer (RLPP). RLPP leverages Monte Carlo (MC) estimators with stochastic sampling operations to generate sufficiently-reliable priors by performing multiple inferences on the foundational medical image segmentation model, MedSAM. Additionally, instead of directly incorporating the priors in the spatial domain, we decompose the cross-attention (CA) mechanism into real symmetric and imaginary anti-symmetric parts via fast Fourier transform (FFT), resulting in the design of the Guided Frequency Cross-Attention (GFCA) solver. By leveraging the conjugated symmetric property of FFT, GFCA reduces the computational complexity of naive CA by nearly half. Extensive experimental results in various tasks demonstrate the superiority of the proposed LRformer in both effectiveness and efficiency. Kecheng Chen, Jiaxin Huang 0006, Yazhou Ren 0001, Xiaorong Pu |
ACM Multimedia | 7 |
| 2025 | ChronoSelect: Robust Learning with Noisy Labels via Dynamics Temporal MemoryabstractTraining deep neural networks on real-world datasets is often hampered by the presence of noisy labels, which can be memorized by over-parameterized models, leading to significant degradation in generalization performance. While existing methods for learning with noisy labels (LNL) have made considerable progress, they fundamentally suffer from static snapshot evaluations and fail to leverage the rich temporal dynamics of learning evolution. In this paper, we propose ChronoSelect (chrono denoting its temporal nature), a novel framework featuring an innovative four-stage memory architecture that compresses prediction history into compact temporal distributions. Our unique sliding update mechanism with controlled decay maintains only four dynamic memory units per sample, progressively emphasizing recent patterns while retaining essential historical knowledge. This enables precise three-way sample partitioning into clean, boundary, and noisy subsets through temporal trajectory analysis and dual-branch consistency. Theoretical guarantees prove the mechanism’s convergence and stability under noisy conditions. Extensive experiments demonstrate ChronoSelect’s state-of-the-art performance across synthetic and real-world benchmarks. Xiaorong Pu, Yazhou Ren 0001 |
MMAsia | 4 |
| 2025 | Self-Supervised MRI Reconstruction using Weighted SSDU via Dual-Branch Latent DiffusionabstractMagnetic Resonance Imaging (MRI) plays a critical role in clinical diagnostics, but its broader application is constrained by prolonged acquisition times, especially under high acceleration factors. We therefore propose a self-supervised MRI reconstruction framework that integrates latent diffusion models with a parallel dual-network architecture, aiming to achieving high-quality image reconstruction under highly accelerated conditions. Within this framework, a dual-branch network is designed to process distinct subsets of under-sampled k-space data. This design enables the capture of complementary information across subsets while effectively mitigating overfitting. Both branches are guided by reconstruction and difference losses to enforce consistency across unobserved regions, enabling accurate recovery of fully sampled images. A solid mathematical foundation is formulated to offer theoretical guarantees regarding the reliable approximation of fully sampled data under specific constraints. Experimental results demonstrate that integrating parallel dual-network selfsupervision with latent diffusion advances self-supervised MRI reconstruction, highlighting its potential for clinically feasible high-fidelity MRI under high acceleration. Jiakai Li, Xiaorong Pu |
SMC | 4 |
| 2025 | Multi-modal isolated sign language recognition based on self-paced learning
Yazhou Ren 0001, Jingyu Pu, Xiaorong Pu, Siyuan Jing, Lifang He 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Variational Graph Generator for Multiview Graph ClusteringabstractMultiview graph clustering (MGC) methods are increasingly being studied due to the explosion of multiview data with graph structural information. The critical point of MGC is to better utilize view-specific and view-common information in features and graphs of multiple views. However, existing works have an inherent limitation that they are unable to concurrently utilize the consensus graph information across multiple graphs and the view-specific feature information. To address this issue, we propose a variational graph generator for MGC (VGMGC). Specifically, a novel variational graph generator is proposed to extract common information among multiple graphs. This generator infers a reliable variational consensus graph based on a priori assumption over multiple graphs. Then, a simple yet effective graph encoder in conjunction with the multiview clustering objective is presented to learn the desired graph embeddings for clustering, which embeds the inferred view-common graph and view-specific graphs together with features. Finally, theoretical results illustrate the rationality of the VGMGC by analyzing the uncertainty of the inferred consensus graph with the information bottleneck (IB) principle. Extensive experiments demonstrate the superior performance of our VGMGC over state-of-the-art methods (SOTAs). The source code is publicly available at: https://github.com/cjpcool/VGMGC. Jianpeng Chen, Yawen Ling, Jie Xu 0044, Yazhou Ren 0001, Shudong Huang, Xiaorong Pu, Zhifeng Hao 0004, Philip S. Yu, Lifang He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Deep Clustering: A Comprehensive SurveyabstractCluster analysis plays an indispensable role in machine learning and data mining. Learning a good data representation is crucial for clustering algorithms. Recently, deep clustering (DC), which can learn clustering-friendly representations using deep neural networks (DNNs), has been broadly applied in a wide range of clustering tasks. Existing surveys for DC mainly focus on the single-view fields and the network architectures, ignoring the complex application scenarios of clustering. To address this issue, in this article, we provide a comprehensive survey for DC in views of data sources. With different data sources, we systematically distinguish the clustering methods in terms of methodology, prior knowledge, and architecture. Concretely, DC methods are introduced according to four categories, i.e., traditional single-view DC, semi-supervised DC, deep multiview clustering (MVC), and deep transfer clustering. Finally, we discuss the open challenges and potential future opportunities in different fields of DC. Yazhou Ren 0001, Jingyu Pu, Zhimeng Yang, Jie Xu 0044, Guofeng Li, Xiaorong Pu, Philip S. Yu, Lifang He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Sparse Bayesian Deep Learning for Cross Domain Medical Image ReconstructionabstractCross domain medical image reconstruction aims to address the issue that deep learning models trained solely on one source dataset might not generalize effectively to unseen target datasets from different hospitals. Some recent methods achieve satisfactory reconstruction performance, but often at the expense of extensive parameters and time consumption. To strike a balance between cross domain image reconstruction quality and model computational efficiency, we propose a lightweight sparse Bayesian deep learning method. Notably, we apply a fixed-form variational Bayes (FFVB) approach to quantify pixel-wise uncertainty priors derived from degradation distribution of the source domain. Furthermore, by integrating the uncertainty prior into the posterior sampled through stochastic gradient Langevin dynamics (SGLD), we develop a training strategy that dynamically generates and optimizes the prior distribution on the network weights for each unseen domain. This strategy enhances generalizability and ensures robust reconstruction performance. When evaluated on medical image reconstruction tasks, our proposed approach demonstrates impressive performance across various previously unseen domains. Jiaxin Huang 0006, Yazhou Ren 0001, Aodi Yang, Xiaorong Pu |
AAAI | 7 |
| 2024 | Adaptive Feature Imputation with Latent Graph for Deep Incomplete Multi-View ClusteringabstractIn recent years, incomplete multi-view clustering (IMVC), which studies the challenging multi-view clustering problem on missing views, has received growing research interests. Previous IMVC methods suffer from the following issues: (1) the inaccurate imputation for missing data, which leads to suboptimal clustering performance, and (2) most existing IMVC models merely consider the explicit presence of graph structure in data, ignoring the fact that latent graphs of different views also provide valuable information for the clustering task. To overcome such challenges, we present a novel method, termed Adaptive feature imputation with latent graph for incomplete multi-view clustering (AGDIMC). Specifically, it captures the embbedded features of each view by incorporating the view-specific deep encoders. Then, we construct partial latent graphs on complete data, which can consolidate the intrinsic relationships within each view while preserving the topological information. With the aim of estimating the missing sample based on the available information, we utilize an adaptive imputation layer to impute the embedded feature of missing data by using cross-view soft cluster assignments and global cluster centroids. As the imputation progresses, the portion of complete data increases, contributing to enhancing the discriminative information contained in global pseudo-labels. Meanwhile, to alleviate the negative impact caused by inferior impute samples and the discrepancy of cluster structures, we further design an adaptive imputation strategy based on the global pseudo-label and the local cluster assignment. Experimental results on multiple real-world datasets demonstrate the effectiveness of our method over existing approaches. Jingyu Pu, Chenhang Cui, Xinyue Chen 0004, Yazhou Ren 0001, Xiaorong Pu, Zhifeng Hao 0005, Philip S. Yu, Lifang He 0001 |
AAAI | 5 |
| 2024 | Homophily-Related: Adaptive Hybrid Graph Filter for Multi-View Graph ClusteringabstractRecently there is a growing focus on graph data, and multi-view graph clustering has become a popular area of research interest. Most of the existing methods are only applicable to homophilous graphs, yet the extensive real-world graph data can hardly fulfill the homophily assumption, where the connected nodes tend to belong to the same class. Several studies have pointed out that the poor performance on heterophilous graphs is actually due to the fact that conventional graph neural networks (GNNs), which are essentially low-pass filters, discard information other than the low-frequency information on the graph. Nevertheless, on certain graphs, particularly heterophilous ones, neglecting high-frequency information and focusing solely on low-frequency information impedes the learning of node representations. To break this limitation, our motivation is to perform graph filtering that is closely related to the homophily degree of the given graph, with the aim of fully leveraging both low-frequency and high-frequency signals to learn distinguishable node embedding. In this work, we propose Adaptive Hybrid Graph Filter for Multi-View Graph Clustering (AHGFC). Specifically, a graph joint process and graph joint aggregation matrix are first designed by using the intrinsic node features and adjacency relationship, which makes the low and high-frequency signals on the graph more distinguishable. Then we design an adaptive hybrid graph filter that is related to the homophily degree, which learns the node embedding based on the graph joint aggregation matrix. After that, the node embedding of each view is weighted and fused into a consensus embedding for the downstream task. Experimental results show that our proposed model performs well on six datasets containing homophilous and heterophilous graphs. Zichen Wen, Yawen Ling, Yazhou Ren 0001, Jianpeng Chen, Xiaorong Pu, Lifang He 0001 |
AAAI | 6 |
| 2024 | Frequency-Constraint VQ-VAE for Adaptive MRI Segmentation
Kecheng Chen, Yazhou Ren 0001, Xiaorong Pu |
ICONIP (9) | 7 |
| 2024 | Dynamic Weighted Graph Fusion for Deep Multi-View Clustering
Yazhou Ren 0001, Jingyu Pu, Chenhang Cui, Xinyue Chen 0004, Xiaorong Pu, Lifang He 0001 |
IJCAI | 6 |
| 2024 | Integrating Vision-Language Semantic Graphs in Multi-View Clustering
Junlong Ke, Zichen Wen, Yechenhao Yang, Chenhang Cui, Yazhou Ren 0001, Xiaorong Pu, Lifang He 0001 |
IJCAI | 6 |
| 2024 | Cross-View Contrastive Fusion for Enhanced Molecular Property Prediction
Yazhou Ren 0001, Jing He 0004, Xiaorong Pu, Lifang He 0001 |
IJCAI | 6 |
| 2024 | Cross-view Contrastive Unification Guides Generative Pretraining for Molecular Property PredictionabstractMulti-view based molecular properties prediction learning has received widely attention in recent years in terms of its potential for the downstream tasks in the field of drug discovery. However, the consistency of different molecular view representations and the full utilization of complementary information among them in existing multi-view molecular property prediction methods remain to be further explored. Furthermore, most current methods focus on generating global level representations at the graph level with information from different molecular views (e.g., 2D and 3D views) assuming that the information can be corresponded to each other. In fact it is not unusual that for example the conformation change or computational errors may lead to discrepancies between views. To addressing these issues, we propose a new Cross-View contrastive unification guides Generative Molcular pre-trained model, call MolCVG. We first focus on common and private information extraction from 2D graph views and 3D geometric views of molecules, Minimizing the impact of noise in private information on subsequent strategies. To exploit both types of information in a more refined way, we propose a cross-view contrastive unification strategy to learn cross-view global information and guide the reconstruction of masked nodes, thus effectively optimizing global features and local descriptions. Extensive experiments on real-world molecular data sets demonstrate the effectiveness of our approach for molecular property prediction task. Xinyue Chen 0004, Yazhou Ren 0001, Xiaorong Pu, Jing He 0004 |
ACM Multimedia | 5 |
| 2024 | Dual-Optimized Adaptive Graph Reconstruction for Multi-View Graph ClusteringabstractMulti-view clustering is an important machine learning task for multi-media data, encompassing various domains such as images, videos, and texts. Moreover, with the growing abundance of graph data, the significance of multi-view graph clustering (MVGC) has become evident. Most existing methods focus on graph neural networks (GNNs) to extract information from both graph structure and feature data to learn distinguishable node representations. However, traditional GNNs are designed with the assumption of homophilous graphs, making them unsuitable for widely prevalent heterophilous graphs. Several techniques have been introduced to enhance GNNs for heterophilous graphs. While these methods partially mitigate the heterophilous graph issue, they often neglect the advantages of traditional GNNs, such as their simplicity, interpretability, and efficiency. In this paper, we propose a novel multi-view graph clustering method based on dual-optimized adaptive graph reconstruction, named DOAGC. It mainly aims to reconstruct the graph structure adapted to traditional GNNs to deal with heterophilous graph issues while maintaining the advantages of traditional GNNs. Specifically, we first develop an adaptive graph reconstruction mechanism that accounts for node correlation and original structural information. To further optimize the reconstruction graph, we design a dual optimization strategy and demonstrate the feasibility of our optimization strategy through mutual information theory. Numerous experiments demonstrate that DOAGC effectively mitigates the heterophilous graph problem. Zichen Wen, Yazhou Ren 0001, Yawen Ling, Chenhang Cui, Xiaorong Pu, Lifang He 0001 |
ACM Multimedia | 6 |
| 2024 | Cross-View Mutual Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation has gained increasing attention due to its potential to alleviate the manual annotation burden. Mainstream methods typically involve two subnets, and conduct a consistency objective to ensure them producing consistent predictions for unlabeled data. However, they often ignore that the complementarity of model predictions is equally crucial. To realize the potential of the multi-subnet architecture, we propose a novel cross-view mutual learning method with a two-branch co-training framework. Specifically, we first introduce a novel conflict-based feature learning (CFL) that encourages the two subnets to learn distinct features from the same input. These distinct features are then decoded into complementary model predictions, allowing both subnets to understand the input from different views. More importantly, we propose a cross-view mutual learning (CML) to maximize the effectiveness of CFL. This approach requires only modifications to the model inputs and supervisory signals, and implements a heterogeneous consistency objective to fully explore the complementarity of model predictions. Consequently, the aggregated predictions can effectively capture both consistency and complementarity across two subnets. Experimental results on three public datasets demonstrate the superiority of CML over previous SoTA methods. Code is available at https://github.com/SongwuJob/CML. Xinyue Chen 0004, Yazhou Ren 0001, Jing He 0004, Xiaorong Pu |
ACM Multimedia | 6 |
| 2024 | Bridging Gaps: Federated Multi-View Clustering in Heterogeneous Hybrid ViewsabstractRecently, federated multi-view clustering (FedMVC) has emerged to explore cluster structures in multi-view data distributed on multiple clients. Many existing approaches tend to assume that clients are isomorphic and all of them belong to either single-view clients or multi-view clients. While these methods have succeeded, they may encounter challenges in practical FedMVC scenarios involving heterogeneous hybrid views, where a mixture of single-view and multi-view clients exhibit varying degrees of heterogeneity. In this paper, we propose a novel FedMVC framework, which concurrently addresses two challenges associated with heterogeneous hybrid views, i.e., client gap and view gap. To address the client gap, we design a local-synergistic contrastive learning approach that helps single-view clients and multi-view clients achieve consistency for mitigating heterogeneity among all clients. To address the view gap, we develop a global-specific weighting aggregation method, which encourages global models to learn complementary features from hybrid views. The interplay between local-synergistic contrastive learning and global-specific weighting aggregation mutually enhances the exploration of the data cluster structures distributed on multiple clients. Theoretical analysis and extensive experiments demonstrate that our method can handle the heterogeneous hybrid views in FedMVC and outperforms state-of-the-art methods. Xinyue Chen 0004, Yazhou Ren 0001, Jie Xu 0044, Fangfei Lin, Xiaorong Pu, Yang Yang 0002 |
NeurIPS | 5 |
| 2024 | Cross-Domain Low-Dose CT Image Denoising With Semantic Preservation and Noise AlignmentabstractDeep learning (DL)-based Low-dose CT (LDCT) image denoising methods may face domain shift problem, where data from different domains (i.e., hospitals) may have similar anatomical regions but exhibit different intrinsic noise characteristics. Therefore, we propose a plug-and-play model called Lowand High-frequency Alignment (LHFA) to address this issue by leveraging semantic features and aligning noise distributions of different CT datasets, while maintaining diagnostic image quality and suppressing noise. Specifically, the LHFA model consists of a Low-frequency Alignment (LFA) module that preserves semantic features (i.e., low-frequency components) with fewer perturbations from both domains for reconstruction. Notably, a Highfrequency Alignment (HFA) module is proposed to quantify the discrepancy between noise representations (i.e., high-frequency components) in a latent space mapped by an auto-encoder. Experimental results demonstrate that the LHFA model effectively alleviates the domain shift problem and significantly improves the performance of DL-based methods on cross-domain LDCT image denoising task, outperforming other domain adaptationbased methods. Jiaxin Huang 0006, Kecheng Chen, Yazhou Ren 0001, Xiaorong Pu, Ce Zhu |
IEEE Trans. Multim. | 5 |
| 2024 | Self-Weighted Contrastive Fusion for Deep Multi-View ClusteringabstractMulti-view clustering can explore consensus information from multiple views and has attracted increasing attention in the past two decades. However, existing works face two major challenges: i) how to deal with the conflict between learning view-consensus information and reconstructing inconsistent viewprivate information, and ii) how to mitigate representation degeneration caused by implementing the consistency objective for multi-view data. To address these challenges, we propose a novel framework of self-weighted contrastive fusion for deep multi-view clustering (SCMVC). First, our method establishes a hierarchical feature fusion framework, effectively segregating the consistency objective from the reconstruction objective. Then, multi-view contrastive fusion is implemented via maximizing consistency expression between the view-consensus representation and global representation, fully exploring the view consistency and complementary. More importantly, we propose to measure the discrepancy between pairwise representations, and then introduce a self-weighting method, which adaptively strengthens useful views in feature fusion and weakens unreliable views, to mitigate representation degeneration. Extensive experiments on nine public datasets demonstrate that our proposed method achieves state-of-the-art clustering performance. The code is available athttps://github.com/SongwuJob/SCMVC. Yazhou Ren 0001, Jing He 0004, Xiaorong Pu, Shudong Huang, Zhifeng Hao 0004, Lifang He 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Self-Supervised Graph Attention Networks for Deep Weighted Multi-View ClusteringabstractAs one of the most important research topics in the unsupervised learning field, Multi-View Clustering (MVC) has been widely studied in the past decade and numerous MVC methods have been developed. Among these methods, the recently emerged Graph Neural Networks (GNN) shine a light on modeling both topological structure and node attributes in the form of graphs, to guide unified embedding learning and clustering. However, the effectiveness of existing GNN-based MVC methods is still limited due to the insufficient consideration in utilizing the self-supervised information and graph information, which can be reflected from the following two aspects: 1) most of these models merely use the self-supervised information to guide the feature learning and fail to realize that such information can be also applied in graph learning and sample weighting; 2) the usage of graph information is generally limited to the feature aggregation in these models, yet it also provides valuable evidence in detecting noisy samples. To this end, in this paper we propose Self-Supervised Graph Attention Networks for Deep Weighted Multi-View Clustering (SGDMC), which promotes the performance of GNN-based deep MVC models by making full use of the self-supervised information and graph information. Specifically, a novel attention-allocating approach that considers both the similarity of node attributes and the self-supervised information is developed to comprehensively evaluate the relevance among different nodes. Meanwhile, to alleviate the negative impact caused by noisy samples and the discrepancy of cluster structures, we further design a sample-weighting strategy based on the attention graph as well as the discrepancy between the global pseudo-labels and the local cluster assignment. Experimental results on multiple real-world datasets demonstrate the effectiveness of our method over existing approaches. Zongmo Huang, Yazhou Ren 0001, Xiaorong Pu, Shudong Huang, Zenglin Xu, Lifang He 0001 |
AAAI | 3 |
| 2023 | Dual Label-Guided Graph Refinement for Multi-View Graph ClusteringabstractWith the increase of multi-view graph data, multi-view graph clustering (MVGC) that can discover the hidden clusters without label supervision has attracted growing attention from researchers. Existing MVGC methods are often sensitive to the given graphs, especially influenced by the low quality graphs, i.e., they tend to be limited by the homophily assumption. However, the widespread real-world data hardly satisfy the homophily assumption. This gap limits the performance of existing MVGC methods on low homophilous graphs. To mitigate this limitation, our motivation is to extract high-level view-common information which is used to refine each view's graph, and reduce the influence of non-homophilous edges. To this end, we propose dual label-guided graph refinement for multi-view graph clustering (DuaLGR), to alleviate the vulnerability in facing low homophilous graphs. Specifically, DuaLGR consists of two modules named dual label-guided graph refinement module and graph encoder module. The first module is designed to extract the soft label from node features and graphs, and then learn a refinement matrix. In cooperation with the pseudo label from the second module, these graphs are refined and aggregated adaptively with different orders. Subsequently, a consensus graph can be generated in the guidance of the pseudo label. Finally, the graph encoder module encodes the consensus graph along with node features to produce the high-level pseudo label for iteratively clustering. The experimental results show the superior performance on coping with low homophilous graph data. The source code for DuaLGR is available at https://github.com/YwL-zhufeng/DuaLGR. Yawen Ling, Jianpeng Chen, Yazhou Ren 0001, Xiaorong Pu, Jie Xu 0044, Xiaofeng Zhu 0001, Lifang He 0001 |
AAAI | 4 |
| 2023 | Deep Multi-view Subspace Clustering with Anchor GraphabstractDeep multi-view subspace clustering (DMVSC) has recently attracted increasing attention due to its promising performance. However, existing DMVSC methods still have two issues: (1) they mainly focus on using autoencoders to nonlinearly embed the data, while the embedding may be suboptimal for clustering because the clustering objective is rarely considered in autoencoders, and (2) existing methods typically have a quadratic or even cubic complexity, which makes it challenging to deal with large-scale data. To address these issues, in this paper we propose a novel deep multi-view subspace clustering method with anchor graph (DMCAG). To be specific, DMCAG firstly learns the embedded features for each view independently, which are used to obtain the subspace representations. To significantly reduce the complexity, we construct an anchor graph with small size for each view. Then, spectral clustering is performed on an integrated anchor graph to obtain pseudo-labels. To overcome the negative impact caused by suboptimal embedded features, we use pseudo-labels to refine the embedding process to make it more suitable for the clustering task. Pseudo-labels and embedded features are updated alternately. Furthermore, we design a strategy to keep the consistency of the labels based on contrastive learning to enhance the clustering performance. Empirical studies on real-world datasets show that our method achieves superior clustering performance over other state-of-the-art methods. Chenhang Cui, Yazhou Ren 0001, Jingyu Pu, Xiaorong Pu, Lifang He 0001 |
IJCAI | 4 |
| 2023 | Federated Deep Multi-View Clustering with Global Self-SupervisionabstractFederated multi-view clustering has the potential to learn a global clustering model from data distributed across multiple devices. In this setting, label information is unknown and data privacy must be preserved, leading to two major challenges. First, views on different clients often have feature heterogeneity, and mining their complementary cluster information is not trivial. Second, the storage and usage of data from multiple clients in a distributed environment can lead to incompleteness of multi-view data. To address these challenges, we propose a novel federated deep multi-view clustering method that can mine complementary cluster structures from multiple clients, while dealing with data incompleteness and privacy concerns. Specifically, in the server environment, we propose sample alignment and data extension techniques to explore the complementary cluster structures of multiple views. The server then distributes global prototypes and global pseudo-labels to each client as global self-supervised information. In the client environment, multiple clients use the global self-supervised information and deep autoencoders to learn view-specific cluster assignments and embedded features, which are then uploaded to the server for refining the global self-supervised information. Finally, the results of our extensive experiments demonstrate that our proposed method exhibits superior performance in addressing the challenges of incomplete multi-view data in distributed environments. Xinyue Chen 0004, Jie Xu 0044, Yazhou Ren 0001, Xiaorong Pu, Ce Zhu, Xiaofeng Zhu 0001, Zhifeng Hao 0005, Lifang He 0001 |
ACM Multimedia | 4 |
| 2023 | Generative Neutral Features-Disentangled Learning for Facial Expression RecognitionabstractFacial expression recognition (FER) plays a critical role in human-computer interaction and affective computing. Traditional FER methods typically rely on comparing the difference between an examined facial expression and a neutral face of the same person to extract the motion of facial features and filter out expression-irrelevant information. With the extensive use of deep learning, the performance of FER has been further improved. However, existing deep learning-based methods rarely utilize neutral faces. To address this gap, we propose a novel deep learning-based FER method called Generative Neutral Features-Disentangled Learning (GNDL), which draws inspiration from the facial feature manifold. Our approach integrates a neutral feature generator (NFG) that generates neutral features in scenarios where the neutral face of the same subject is not available. The NFG uses fine-grained features from examined images as input and produces corresponding neutral features with the same identity. We train the NFG using a neutral feature reconstruction loss to ensure that the generative neutral features are consistent with the actual neutral features. We then disentangle the generative neutral features from the examined features to remove disturbance features and generate an expression deviation embedding for classification. Extensitive experimental results on three popular databases (CK+, Oulu-CASIA, and MMI) demonstrate that our proposed GNDL method outperforms state-of-the-art FER methods. Zhenqian Wu, Yazhou Ren 0001, Xiaorong Pu, Zhifeng Hao 0005, Lifang He 0001 |
ACM Multimedia | 3 |
| 2023 | A Novel Approach for Effective Multi-View Clustering with Information-Theoretic PerspectiveabstractMulti-view clustering (MVC) is a popular technique for improving clustering performance using various data sources. However, existing methods primarily focus on acquiring consistent information while often neglecting the issue of redundancy across multiple views.
This study presents a new approach called Sufficient Multi-View Clustering (SUMVC) that examines the multi-view clustering framework from an information-theoretic standpoint. Our proposed method consists of two parts. Firstly, we develop a simple and reliable multi-view clustering method SCMVC (simple consistent multi-view clustering) that employs variational analysis to generate consistent information. Secondly, we propose a sufficient representation lower bound to enhance consistent information and minimise unnecessary information among views. The proposed SUMVC method offers a promising solution to the problem of multi-view clustering and provides a new perspective for analyzing multi-view data.
To verify the effectiveness of our model, we conducted a theoretical analysis based on the Bayes Error Rate, and experiments on multiple multi-view datasets demonstrate the superior performance of SUMVC. Chenhang Cui, Yazhou Ren 0001, Jingyu Pu, Xiaorong Pu, Yutao Shi, Lifang He 0001 |
NeurIPS | 5 |
| 2023 | DC-FUDA: Improving deep clustering via fully unsupervised domain adaptation
Zhimeng Yang, Yazhou Ren 0001, Zirui Wu, Ming Zeng 0009, Jie Xu 0044, Yang Yang 0002, Xiaorong Pu, Philip S. Yu, Lifang He 0001 |
Neurocomputing | 7 |
| 2023 | Self-Supervised Discriminative Feature Learning for Deep Multi-View ClusteringabstractMulti-view clustering is an important research topic due to its capability to utilize complementary information from multiple views. However, there are few methods to consider the negative impact caused by certain views with unclear clustering structures, resulting in poor multi-view clustering performance. To address this drawback, we proposeself-supervised discriminative feature learning fordeepmulti-viewclustering (SDMVC). Concretely, deep autoencoders are applied to learn embedded features for each view independently. To leverage the multi-view complementary information, we concatenate all views’ embedded features to form the global features, which can overcome the negative impact of some views’ unclear clustering structures. In a self-supervised manner, pseudo-labels are obtained to build a unified target distribution to perform multi-view discriminative feature learning. During this process, global discriminative information can be mined to supervise all views to learn more discriminative features, which in turn are used to update the target distribution. Besides, this unified target distribution can make SDMVC learn consistent cluster assignments, which accomplishes the clustering consistency of multiple views while preserving their features’ diversity. Experiments on various types of multi-view datasets show that SDMVC outperforms 14 competitors including classic and state-of-the-art methods. The code is available athttps://github.com/SubmissionsIn/SDMVC. Jie Xu 0044, Yazhou Ren 0001, Huayi Tang, Zhimeng Yang, Lili Pan 0001, Yang Yang 0002, Xiaorong Pu, Philip S. Yu, Lifang He 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Cross Domain Low-Dose CT Image Denoising With Semantic Information AlignmentabstractRecently, cross domain adaptation has been applied into quite a few image restoration tasks. While promising performance has been achieved, the domain shift problem between the training set (a.k.a., source domain) and the testing set (a.k.a., target domain) in Low-dose Computed Tomography (LDCT) image denoising tasks is typically ignored by most existing methods. This is prone to the degradation of the denoising performance due to large discrepancy of feature distribution in each dataset from various vendors. Therefore, a simple yet effective LDCT denoising approach has been proposed in this paper to alleviate the domain shift between source and target domains through a novel semantic information alignment. Specifically, we first propose an adaptive version of random frequency mask (RFM) to extract the shared semantic information of cross domains. Then, we incorporate the mask into the existing denoiser to construct a semantic-information-guided objective. Experiments on synthetic and real datasets show our proposed method achieves impressive performance. Jiaxin Huang 0006, Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001 |
ICIP | 4 |
| 2022 | Shared-Attribute Multi-Graph Clustering with Global Self-Attention
Jianpeng Chen, Zhimeng Yang, Jingyu Pu, Yazhou Ren 0001, Xiaorong Pu, Lifang He 0001 |
ICONIP (1) | 5 |
| 2022 | ClusterUDA: Latent Space Clustering in Unsupervised Domain Adaption for Pulmonary Nodule Detection
Kecheng Chen, Xiaorong Pu, Chao Li 0034, Yazhou Ren 0001 |
ICONIP (6) | 5 |
| 2022 | Self-Paced Label Distribution Learning for In-The-Wild Facial Expression RecognitionabstractLabel distribution learning (LDL) has achieved great progress in facial expression recognition (FER), where the generating label distribution is a key procedure for LDL-based FER. However, many existing researches have shown the common problem with noisy samples in FER, especially on in-the-wild datasets. This issue may lead to generating unreliable label distributions (which can be seen as label noise), and will further negatively affect the FER model. To this end, we propose a play-and-plug method of self-paced label distribution learning (SPLDL) for in-the-wild FER. Specifically, a simple yet efficient label distribution generator is adopted to generate label distributions to guide label distribution learning. We then introduce self-paced learning (SPL) paradigm and develop a novel self-paced label distribution learning strategy, which considers both classification losses and distribution losses. SPLDL first learns easy samples with reliable label distributions and gradually steps to complex ones, effectively suppressing the negative impact introduced by noisy samples and unreliable label distributions. Extensive experiments on in-the-wild FER datasets (\emphi.e., RAF-DB and AffectNet) based on three backbone networks demonstrate the effectiveness of the proposed method. Jianjian Shao, Zhenqian Wu, Yuanyan Luo, Shudong Huang, Xiaorong Pu, Yazhou Ren 0001 |
ACM Multimedia | 5 |
| 2022 | TEMDnet: A Novel Deep Denoising Network for Transient Electromagnetic Signal With Signal-to-Image TransformationabstractThe considerable prospecting depth and accurate subsurface characteristics can be obtained by the transient electromagnetic method (TEM) in geophysics. Nevertheless, the time-domain TEM signal received by the coil is easily disturbed by environmental background noise, artificial noise, and electronic noise of the equipment. Recently, deep neural networks (DNNs) have been used to solve the TEM denoising problem and have achieved better performance than traditional methods. However, the existing denoising method with DNN adopts fully connected neural networks and is therefore not flexible enough to deal with various signal scales. To address these issues, a novel denoising framework with deep convolutional neural networks (CNNs) of transforming the TEM signal denoising task into an image denoising task (namely, TEMDnet) is proposed in this article. Specifically, a novel signal-to-image transformation method is developed first to preserve the structural features of TEM signals. Then, a novel deep CNN-based denoiser is proposed to further perform feature learning, in which the residual learning mechanism is adopted to model the noise estimation image for different signal features. Extensive experiments demonstrate that the proposed framework can achieve much better performance compared with other state-of-the-art approaches on both simulated signals and real-world signals from a landfill leachate treatment plant in Chengdu, Sichuan, China. Models and code are available at https://github.com/tonyckc/TEMDnet_demo. Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001, Hang Qiu 0002, Fanqiang Lin, Saimin Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Multi-VAE: Learning Disentangled View-common and View-peculiar Visual Representations for Multi-view ClusteringabstractMulti-view clustering, a long-standing and important research problem, focuses on mining complementary information from diverse views. However, existing works often fuse multiple views’ representations or handle clustering in a common feature space, which may result in their entanglement especially for visual representations. To address this issue, we present a novel VAE-based multi-view clustering framework (Multi-VAE) by learning disentangled visual representations. Concretely, we define a view-common variable and multiple view-peculiar variables in the generative model. The prior of view-common variable obeys approximately discrete Gumbel Softmax distribution, which is introduced to extract the common cluster factor of multiple views. Meanwhile, the prior of view-peculiar variable follows continuous Gaussian distribution, which is used to represent each view’s peculiar visual factors. By controlling the mutual information capacity to disentangle the view-common and view-peculiar representations, continuous visual information of multiple views can be separated so that their common discrete cluster information can be effectively mined. Experimental results demonstrate that Multi-VAE enjoys the disentangled and explainable visual representations, while obtaining superior clustering performance compared with state-of-the-art methods. Jie Xu 0044, Yazhou Ren 0001, Huayi Tang, Xiaorong Pu, Xiaofeng Zhu 0001, Ming Zeng 0009, Lifang He 0001 |
ICCV | 4 |
| 2021 | Lesion-Inspired Denoising Network: Connecting Medical Image Denoising and Lesion DetectionabstractDeep learning has achieved notable performance in the denoising task of low-quality medical images and the detection task of lesions, respectively. However, existing low-quality medical image denoising approaches are disconnected from the detection task of lesions. Intuitively, the quality of denoised images will influence the lesion detection accuracy that in turn can be used to affect the denoising performance. To this end, we propose a play-and-plug medical image denoising framework, namely Lesion-Inspired Denoising Network (LIDnet), to collaboratively improve both denoising performance and detection accuracy of denoised medical images. Specifically, we propose to insert the feedback of downstream detection task into existing denoising framework by jointly learning a multi-loss objective. Instead of using perceptual loss calculated on the entire feature map, a novel region-of-interest (ROI) perceptual loss induced by the lesion detection task is proposed to further connect these two tasks. To achieve better optimization for overall framework, we propose a customized collaborative training strategy for LIDnet. On consideration of clinical usability and imaging characteristics, three low-dose CT images datasets are used to evaluate the effectiveness of the proposed LIDnet. Experiments show that, by equipping with LIDnet, both of the denoising and lesion detection performance of baseline methods can be significantly improved. Kecheng Chen, Kun Long, Yazhou Ren 0001, Xiaorong Pu |
ACM Multimedia | 5 |
| 2021 | Non-Linear Fusion for Self-Paced Multi-View ClusteringabstractWith the advance of the multi-media and multi-modal data, multi-view clustering (MVC) has drawn increasing attentions recently. In this field, one of the most crucial challenges is that the characteristics and qualities of different views usually vary extensively. Therefore, it is essential for MVC methods to find an effective approach that handles the diversity of multiple views appropriately. To this end, a series of MVC methods focusing on how to integrate the loss from each view have been proposed in the past few years. Among these methods, the mainstream idea is assigning weights to each view and then combining them linearly. In this paper, inspired by the effectiveness of non-linear combination in instance learning and the auto-weighted approaches, we propose Non-Linear Fusion for Self-Paced Multi-View Clustering (NSMVC), which is totally different from the the conventional linear-weighting algorithms. In NSMVC, we directly assign different exponents to different views according to their qualities. By this way, the negative impact from the corrupt views can be significantly reduced. Meanwhile, to address the non-convex issue of the MVC model, we further define a novel regularizer-free modality of Self-Paced Learning (SPL), which fits the proposed non-linear model perfectly. Experimental results on various real-world data sets demonstrate the effectiveness of the proposed method. Zongmo Huang, Yazhou Ren 0001, Xiaorong Pu, Lifang He 0001 |
ACM Multimedia | 3 |
| 2021 | Probability-based Mask R-CNN for pulmonary embolism detection
Kun Long, Xiaorong Pu, Yazhou Ren 0001, Mingxiu Zheng, Chunjiang Song, Su Han, Fengbin Deng |
Neurocomputing | 3 |
| 2021 | Dual self-paced multi-view clustering
Zongmo Huang, Yazhou Ren 0001, Xiaorong Pu, Lili Pan 0001, Dezhong Yao 0001, Guoxian Yu |
Neural Networks | 3 |
| 2020 | Low-Dose CT Image Blind Denoising with Graph Convolutional Networks
Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001, Hang Qiu 0002, Haoliang Li |
ICONIP (1) | 2 |
| 2020 | Multi-graph fusion for multi-view spectral clustering
Zhao Kang 0001, Guoxin Shi, Shudong Huang, Wenyu Chen 0001, Xiaorong Pu, Joey Tianyi Zhou, Zenglin Xu |
Knowl. Based Syst. | 5 |
| 2019 | Automatic Identification of Alzheimer's Disease and Epilepsy Based on MRIabstractAlzheimer's disease (AD) and epilepsy are both common chronic diseases in neurology. A certain proportion of AD patients have been found to have epilepsy complication. Neuroimaging such as structural magnetic resonance imaging (MRI) has been proved to be useful in assessing the pathology of AD and epilepsy. Computer-aided diagnosis (CAD) on automatical MRI identification can be applied to assist physicians in diagnosing both diseases. In this paper, it is investigated that the performance of identification on AD, AD complicated with Epilepsy, and Epilepsy based on different MRI brain tissues and feature extraction methods. 17 AD patients, 17 AD patients complicated with epilepsy, 15 epilepsy patients, and 10 healthy control subjects from West China Hospital, Sichuan University were studied. Several preprocessing steps were performed for each MRI to obtain gray matter (GM) and white matter (WM) tissue voxels. Principal component analysis (PCA) and partial least squares (PLS) were adopted to extract features. Three classes of patients and healthy controls were distinguished separately by support vector machine (SVM). The performance is evaluated by k-fold cross-validation strategy. The approach on combination of GM and WM tissues with PCA archieved the optimal performance, with the accuracy of 87.41%, 83.7%, and 75.2% for AD, AD complicated with epilepsy, and epilepsy identification respectively. Our proposed approach appears to have significant potential as a clinical decision support tool in assisting physicians in their clinical routine. Xijue Zhang, Wanling Li, Wangshu Shen, Xiaorong Pu |
ICTAI | 5 |
| 2018 | Median local ternary patterns optimized with rotation-invariant uniform-three mapping for noisy texture classification
Luping Ji, Xiaorong Pu, Guisong Liu |
Pattern Recognit. | 3 |
| 2018 | Training-Based Gradient LBP Feature Models for Multiresolution Texture ClassificationabstractLocal binary pattern (LBP) is a simple, yet efficient coding model for extracting texture features. To improve texture classification, this paper designs a median sampling regulation, defines a group of gradient LBP (gLBP) descriptors, proposes a training-based feature model mapping method, and then develops a texture classification frame using the multiresolution feature fusion of four gLBP descriptors. Cooperated by median sampling, four descriptors encode a pixel respectively by central gradient, radial gradient, magnitude gradient and tangent gradient to generate initial gLBP patterns. The feature mapping models of gLBP descriptors are constructed by the maximal relative-variation rate (mr2) of rotation-invariant patterns, and then prestored as mapping lookup files. By mapping, initial patterns can be transformed into low-dimensional ones. And then it generates multiresolution texture features via the joint and concatenation of gLBP descriptors on different sampling parameters. A trained nearest neighbor classifier with chi-square distance is applied to classify textures by feature histograms. The experimental results of simulation on five public texture databases show that the proposed method is reliable and efficient in texture classification. In comparison with nine other similar approaches, including two state-of-the-art ones, the proposed method runs faster than most of them and also outperforms all of them in terms of classification accuracy and noise robustness. It achieves higher accuracy and has also better robustness to the Salt&Pepper and Gaussian noise added artificially into texture images. Luping Ji, Guisong Liu, Xiaorong Pu |
IEEE Trans. Cybern. | 4 |
| 2017 | Deep Semantics-Preserving Hashing Based Skin Lesion Image Retrieval
Xiaorong Pu, Hang Qiu 0002, Yinhui Sun |
ISNN (2) | 1 |
| 2015 | One-dimensional pairwise CNN for the global alignment of two DNA sequences
Luping Ji, Xiaorong Pu, Hong Qu 0002, Guisong Liu |
Neurocomputing | 2 |
| 2015 | Facial expression recognition from image sequences using twofold random forest classifier
Xiaorong Pu, Luping Ji, Zhihu Zhou |
Neurocomputing | 1 |
| 2011 | Chaotic Modeling of Time-Delay Memristive System
Ju Jin, Yongbin Yu 0001, Xiaorong Pu, Xiaofeng Liao 0001 |
ICIC (1) | 4 |
| 2009 | A New Incremental PCA Algorithm With Application to Visual Learning and Recognition
Zhang Yi 0001, Xiaorong Pu |
Neural Process. Lett. | 3 |
| 2009 | Manifold-Based Learning and SynthesisabstractThis paper proposes a new approach to analyze high-dimensional data set using low-dimensional manifold. This manifold-based approach provides a unified formulation for both learning from and synthesis back to the input space. The manifold learning method desires to solve two problems in many existing algorithms. The first problem is the local manifold distortion caused by the cost averaging of the global cost optimization during the manifold learning. The second problem results from the unit variance constraint generally used in those spectral embedding methods where global metric information is lost. For the out-of-sample data points, the proposed approach gives simple solutions to transverse between the input space and the feature space. In addition, this method can be used to estimate the underlying dimension and is robust to the number of neighbors. Experiments on both low-dimensional data and real image data are performed to illustrate the theory. Zhang Yi 0001, Xiaorong Pu |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2008 | A new local PCA-SOM algorithm
Zhang Yi 0001, Xiaorong Pu |
Neurocomputing | 3 |
| 2008 | Holistic and partial facial features fusion by binary particle swarm optimization
Xiaorong Pu, Zhang Yi 0001, Zhongjie Fang |
Neural Comput. Appl. | 1 |
| 2007 | Multilayer Perceptron Networks Training Using Particle Swarm Optimization with Minimum Velocity Constraints
Xiaorong Pu, Zhongjie Fang, Yongguo Liu |
ISNN (3) | 1 |
| 2007 | Binary Fingerprint Image Thinning Using Template-Based PCNNsabstractThis correspondence presents a coarse-to-fine binary-image-thinning algorithm by proposing a template-based pulse-coupled neural-network model. Under the control of coupled templates, this algorithm iteratively skeletonizes a binary image by changing the load signals of pulse neurons. A direction-constraining scheme for avoiding fingerprint ridge spikes has been discussed. Experiments show that this algorithm is effective for fingerprint thinning, as well as other common images. Moreover, this algorithm can be coupled with a fingerprint identification system to improve the recognition performance. Luping Ji, Zhang Yi 0001, Lifeng Shang, Xiaorong Pu |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2006 | Recognizing Partially Damaged Facial Images by Subspace Auto-associative Memories
Xiaorong Pu, Zhang Yi 0001 |
ISNN (2) | 1 |
| 2006 | Parts-Based Holistic Face Recognition with RBF Neural Networks
Xiaorong Pu, Ziming Zheng |
ISNN (2) | 2 |
| 2005 | Face Recognition Using Fisher Non-negative Matrix Factorization with Sparseness Constraints
Xiaorong Pu, Zhang Yi 0001, Ziming Zheng, Mao Ye 0001 |
ISNN (2) | 1 |